Everything is both legal and illegal until a lawsuit happens. Then it collapses depending a little bit on the facts and mostly on who has the better lawyers.
I suspect Nitter's first round with lawyers pointed out that scraping is legal, but now they have been threatened with something else than scraping - Elon claims something else the way Nitter runs is illegal, such as the use of fake accounts to circumvent an access control device (DMCA 1201).
At the end of the day, as an individual or a team or a company, regardless of the statue and case law, you have to perform the calculus on your monetary and legal resources versus your counterparty.
Obviously ungrounded and frivolous cases tend to be easier to defend asymmetrically, but if I was X's legal team, there's no shortage of semi-plauisble claims I could throw at the wall and see what sticks.
As an example of this imbalance in action, BrightData is a 'gray area' company that basically does this exact kind of scraping. They have somehow won against Meta Platforms suing them, and even got X's lawsuit against them for scraping -- identical (?) activity to XCancel -- dismissed.
According to Wikipedia:
> In May 2024, a federal judge dismissed the suit, ruling that Bright Data did not violate X's terms of service or copyright by scraping publicly accessible data.[24] The judge emphasized that such scraping practices are generally legal and that restricting them could lead to information monopolies
But does XCancel have the resources of a company like Bright Data, that's funded and used by companies like Deloitte and Moodys?
If the name Bright Data is ringing a bell to anyone, it’s probably because they are a (the?) primary offender running the LG TV “residential proxy” (botnet)
Which, btw, was the least bad thing about those TVs. I support residential proxying as a means to liberate public data that's being held captive by people like Elon.
Notice it says "terms of service or copyright". If X's lawyers have any intelligence, they'll have a reason why XCancel is not identical to Bright Data. Perhaps this time, instead of claiming it's a copyright violation, they'll claim it's wire fraud because multiple accounts are used.
Anyone can make an account, though, and instantly access that content, so Elon doesn't really have any leg to stand on by claiming they're private. It's not the same as Cambridge Analytica scraping stuff you had to have certain privileges to see by tricking the system into granting those privileges.
Elon unfortunately may have a leg to stand on, because it doesn't functionally matter whether "anyone can make an account and see it". I don't believe any legal ruling to date has been fine with that distinction.
It depends what you mean by "not allowed". Breaking ToS isn't a crime, but you might be sued for damages, but how can they show there were any damages?
But if they ban the nitter operator surely you can't just bypass that ban legally by making a new X account. That there are bans and the operator has to bypass bans by pretending to be a new person I think that this makes it different.
It's like a club that checks ids to ban people. I wouldn't call that club open to the general public.
Tall order. The main Nitter instances were using tens of thousands of accounts to evade detection, IIRC. Adding money to the mix creates more legal liability and a paper trail towards people who can be sued.
It's much easier if thousands of people set up their own Nitter instances.
It's not that cut and dry or else search engines wouldn't be legal. It depends on how much is used, for what context, etc. This very well may wind up being fair use.
Search engines "modify" it ie. show snippets + direct to the actual site.
In general fair use pretty much always requires it to be transformative and/or point to the source. Simply scraping it to prevent people from going to X isn't free use in any definition I've heard.
I wonder what legal gymnastics are needed for "I can scrape anything off the web ignoring copyright and build a product from this, but you can’t even display what’s on my webpage elsewhere”.
Perhaps that is it in fact. The act of protecting it from scraping means you object. 99% of the blogged contents etc. Big AI helped themselves to was just… there. Public. Not free from copyright but still not paywalled.
Precedent is pretty clear: competitive uses bad, transformative uses good. Xcancel scrapes and then competes directly with X, whereas LLM labs scrape the internet to make an agentic intelligent bot, a transformative use of the scraped content.
> Xcancel scrapes and then competes directly with X, whereas LLM labs scrape the internet to make an agentic intelligent bot, a transformative use of the scraped content.
It seems unreasonable to stop there though; the agentic bots are designed and marketed as able to compete with the initially-scraped sources.
I'm not convinced that a competitive use at one remove should be treated as not competitive.
I think that's more true in image generation than in text? At least, all the money is in LLMs that write code, not LLMs that write O'Reilley-style coding books.
If you have a websites that offers guides, how-tos or tutorials, LLMs directly compete with you. StackOverflow would also have a really good case
After all LLMs don't just code, they also answer questions and give step-by-step instructions. In terms of total userbase those features are used a lot more than writing code
If a society operates under a rule like this, it is no better than Russia or any other tyranny where law is for me but not for thee. This is not how it should work in a supposedly free and lawful country.
It’s been working like that in the US for a while.
Do you think the bottom 99% of the country would ever win a legal fight with one of the tech billionaires?
Even if they were 100% in the right, they could just drag out the legal process with endless motions and appeals until the average Joe lacked the funds to continue the fight.
There's a difference between creating a market for something better, so that nobody wants the old thing, and competing _in_ the market for the old thing by copying it directly.
It's like saying "toaster oven/air fryer combos" don't actually compete with toaster ovens or air fryers because they are creating a market for something better
Of course they complete.
Toaster ovens compete with toasters. Microwaves compete with toaster ovens.
Just because it's not the exact same product doesn't mean it's not competing
Would I be allowed to steal LG's designs for a microwave and make a "superwave" that does laundry and heats food? Would you claim those products don't compete because the superwave is "something better"?
If LLMs only made SO redundant by writing code and solving my technical problems autonomously so I never have to think about it, I would agree. But often I do ask LLMs technical questions, and they answer in great detail. And that part is a very direct SO competitor
And what would be a read-only version of X like XCancel compete against, exactly? Ads impressions? That would be the only possible thing yet they don't add any ads.
It's depriving X of impressions that they could monetise, no? Xcancel doesn't have to make money itself, it just has to impair the rights of the copyright holder. Otherwise piracy would also be legal as long as it were non-profit...
> Otherwise piracy would also be legal as long as it were non-profit...
Which is in a few jurisdictions, or at least is not prosecuted if it's for personal use.
Also, according to your definition, the creator of uBlock Origin or any other adblock system should be sued in the same way, because they are depriving $ADS_CORP of their precious impressions.
Well, adblockers don't copy the copyrighted content. They just control how it's rendered on the user's machine. Copyright cares about making copies and especially distributing them.
Xcancel isn't copying the copyrighted content either. It takes the raw JSON/HTML data from twitter/x via reverse proxy and redisplays it on xcancel, as opposed to twitter/x. How is that copying?
You have a point on this, I recognize, but it still seems a very thin line to walk (for X) - at least morally, because I don't think they are actually loosing real big money to anyone.
X doesnt own msot of those copyrights, except where its elon musk's own posts.
i dont think the actual copyright owners care, given they put their content onto a vaguely public view where they aim to get the most traffic to something else they are doing
Bartz v Anthropic. Though the plaintiffs did get something, it was because of the piracy to the original works (competing against the legal market for the books), not the use of them to train the LLM.
That would definitely limit damages to strictly those accounts.
I'm not even sure he can use his own account as one of them. The SEC might be pretty friendly to him but I'm not sure that limiting access to a location where material information about Tesla/SpaceX is provided won't become a problem.
But I'm not even sure what damages the accounts are suffering as revenue sharing is going away [1]. With Bartz the damage is a loss of sale. With X the damage is $0 per post to the poster.
There is a newer Original Content Rewards program [2] but it seems to split revenue from X Premium and presumably people that have X Premium are not using XCancel so the damages would be 0.
How clear is it when I google a recipe and get an AI-generated recipe that's clearly derived from the top three results and then placed above those results? That sounds like it's both transformative (in that the recipe created by the AI may not match any one of the scraped recipes perfectly) and also competitive (in that the AI takes page views away from the pages it got the recipes from)
fair, but recipes aren't the only thing where AI summaries at the top of search results are simultaneously transformative and competitive. Any information that is scraped from a website and then summarized by AI in a way that prevents that website from getting views is both transformative and competitive.
Taking something someone else made and showing it as-is, bypassing their own restrictions: No no.
Taking something someone else made, modify it or use parts of it in some bigger thing or completely change it: Fine, if you have money and/or run a company
So in theory if you took twitter content and then transformed it so it "summarizes" all tweets with an AI rather than posting the exact text, would that be allowed?
The point is that you can't steal someone else's content 1:1. But you can use it for a different use (say, display the tweet in an article, then comment on it).
That's exactly the way UK courts are heading, see Getty vs Stability AI. The court ruled that there's no infringment because the model doesn't store exact copies, just derived weights, and therefore when it generates new images those aren't copies of protected works.
It should have nothing to do with storing copies it should have to do with what the models can produce. And it's clear they can produce copyrighted works, they've just been tuned so they don't.
They can't unless the tuner produces a copyrighted work for whatever their purpose would be, because they can step in and "detune" the thing at 3am for a competitive advantage.
You can produce copyrighted works and have just been tuned not to lol. I fail to see how limiting an ability to comply with the law is any different from just complying with the law?
scraping content is mostly legal, redistributing content is not.
if you started doing the same to, say, instagram content both meta and individual creators would sue you as well.
sites like archive.ph are in a similar bucket btw, and yet nobody's complaining (except websites seeing people evading their paywall). but at the end of the day it's not really fair to apply laws differentially on the basis of whose political ideas we like more.
Meta has no exclusive rights to the content on Instagram, and X has no exclusive rights to the content on X. They have a non-exclusive license to republish it, etc.
I love how the content belongs to them when someone else reposts it but it belongs to the user if the content is illegal. Such a double standard with these social media and AI companies. Why do we put up with it?
You let it happen. Once people stop letting it happen, it'll stop. But social media is apparently the new "opium of the masses" so here we are and no one wants to do anything.
I am also doing something, by hosting one of the public instances of Nitter.
Nitter is useful for sporadic random access to tweets, but for public feeds like municipal authorities etc. it would be useful if someone scraped the feed and re-hosted the feed from their own server, without being hobbled by rate limits. Is that what you're doing?
Yes, exactly. I limit the scraping to popular enough accounts. The wikimedia community is already organized to decide who deserves a Wikipedia page so I use that.
We put up with it mostly because it enables a lot of communication to happen at all.
You want to hold the corporations legally liable for the content their host, you can say goodbye to basically reddit as a whole, any twitter clone, youtube comment sections and a whole more stuff.
Spot on. This is where a lot of these "terms and conditions" break down logically. Viewing some content on the internet is literally copying it.
So is the distinction that xcancel served the content? But when I run
mtr xcancel.com
I see a bunch of hops between me and them. Every one of those hops is literally copying and retransmitting all the content. Are they not also serving it?
No, this is where programmers rules-lawyer in ways that actual lawyers don't and then get law stuff hilariously wrong. No judge thinks that viewing an HTML page is downloading it, because downloading means saving a copy to your computer, not just looking at it. Even having an internet cache folder doesn't count as downloading. Even copying the file from the internet cache folder to somewhere might not count as downloading, although it'd still be a copy.
Same as when LG said their TVs don't record you and then Hacker News said "how can they detect voice commands if they don't record your voice"... facepalm.
It makes more sense when you remember it's not a computer program and the things that are written in the law are not the things that will actually happen in the way that "if(foo) bar;" makes bar happen if foo is true. It's more like a book of excuses you could use for why you didn't do your homework.
Then the other side also has to bring an excuse for why you were supposed to do it, and if the principal thinks their excuse is better than yours, you get detention.
If you tell the principal "I don't have to do my homework because work means employment and it's illegal to employ a minor" you'll get detention for not doing your homework and extra detention for being a smartass.
And this example is not just due to people not taking the trouble to write fully specified rules. I don't think such rules could even be written. You can just do your best to cover the cases you can think of. The complexity of society is incomprehensibly vast and constantly changing, and the law has to have wiggle room to account for it.
You don't want fully-specified rules because a rule with strict boundaries has loopholes. You actually want a clearly allowed area, a clearly disallowed area, and a gradually increasing gradient of punishment in between, so that a small change in behaviour produces only a small change in punishment, and avoiding punishment requires a large change in behaviour.
Could you elaborate in what way you find the law mostly doesn't make sense? It has to be flexible in order to work with actual humans. Why should visiting a page on your computer count as copying? Usually when we talk about copying it's someone making a duplicate so it can be accessed later. Only a very technical user is going to be diving into their cache to view that content after the fact. The vast majority of people don't understand that the browser is storing anything on their computer, much less how to access it before it's purged.
I can't remember the court case, but Blizzard did argue and win in court that WoW Glider's producers violated copyright law. If I recall correctly violating the TOS meant that an unauthorized copy made by executing the file chasing it to load WoW into RAM was created.
It looks like that was MDY Industries, LLC v. Blizzard Entertainment, Inc., which relied on MAI Systems Corp. v. Peak Computer, Inc. for the relevant part of the ruling.
The person I was responding to was saying that anytime you viewed copyrighted content with a browser you’d necessarily be committing copyright infringement. I’m not a lawyer but I can imagine that the reasoning there would be slightly different from someone simply viewing a post in a browser as part of the intended use of the site.
Oh yeah, I understood your point, but given MDY Industries, LLC v .Blizzard who knows what the "right" judge would rule? With IP laws these days we're really getting into weird places.
> Why should visiting a page on your computer count as copying?
Because there's no physical mechanism for the information to be transmitted over a computer network other than by copying the bytes.
Note this is distinct from broadcast systems like analog television or radio. Packet switching networks only function by copying information and storing multiple copies around the internet, including in your computer's RAM (and disk, if cached).
So a legal definition that says "this kind of copying is copying but that other kind of copying isn't copying" makes no sense at all. Like many other legal definitions--it's all about what has been successfully snuck past a jury at one point or another in the past, without any heed for how things actually work.
It's not about "how things actually work", the law is there to regulate human activity. The law tends to call these copies on the wire, in RAM, in caches, etc. "transient copies", which is fine until a human starts using them as non-transient copies, e.g. saves them for later.
You could argue that your MP3 of Enjoy the Silence is actually just a big number, and you can XOR it with 0xFF and it's a completely different big number, and you just happen to XOR it with 0xFF when you want to listen to it. The courts would look past that, and instead determine if you created that "big number" by MP3-encoding the track from a CD you owned (legal), versus obtaining it from some file-sharing network (not legal)
> Because there's no physical mechanism for the information to be transmitted over a computer network other than by copying the bytes.
Your response seems to ignore everything in my comment other than the second sentence. I was asking why that detail should matter as far as the law is concerned, and I gave some reasons I don't think that would be good or practical.
If your link is set up to make the image display immediately (that is, you wrap it in image tags, or as in one case, embed Instagram posts) then you may be violating copyright. What's more, in Europe, just a hyperlink to a copyrighted work violates copyright.
Conclusion: copyright is not about copying, it's about access.
Sure, but that seems different from what I was addressing. The person I was responding to was saying that the law as a whole usually doesn’t make sense. They were saying that in the context of arguing that if the law didn’t consider viewing a page of copyrighted copying as involving copying due to the technical basis of it having to transfer bits to your computer then the law didn’t make sense. My point was that laws don’t have to encompass or fully specify all edge cases, and that the ways laws are written can be open to interpretation. I think I removed a sentence before posting about the purpose of finders of facts in the US system like juries or judges in bench trials.
> Your response seems to ignore everything in my comment other than the second sentence.
I deliberately ignored it, because it was all irrelevant.
> I was asking why that detail should matter as far as the law is concerned, and I gave some reasons I don't think that would be good or practical.
I have no idea at all how anything should matter as far as the law is concerned. Not my problem, unless I somehow get caught. But not getting caught is a problem grounded in reality, unlike legal ones. I think I can manage that.
That said, if laws about computers don't comport with how computers actually work, I'll take extra amounts of glee in violating them.
And, even more gleefully, nobody will be able to detect my violations. My internet traffic will look identically the same as someone "innocently copying" or whatever.
> If you serve as a mere conduit for automatic transmission of user communications, there are no other qualifications or obligations you need to meet. If you serve a caching function, in addition to the two requirements above, you must maintain comply with the notice-and-takedown process.
Distilling isn’t copying and redistributing, for the same reason that you reading a story and then writing your own story based on ideas you learned is different from you reading a book, writing all the words down verbatim, and then publishing it as your own.
That is what all LLMs could do in 2023, verbatim, before they trained it out of them in order to keep up the pretense that there is no plagiarism. Now they all obfuscate the original or refuse to cite.
reply