Hi Crowd!
Search is cooked. Which might mean the internet as we know is about to very different. That’s a bold claim, but hear me out. If you are not someone who makes your livelihood by posting things on the internet it may come as a surprise to you that publishers have a somewhat complicated relationship with search (Google specifically, but all search in general). Pull a chair up around the fire and I’ll tell you a story.
In the beginning it was good, if someone wanted to learn something they could search for it and if you had a site about that thing the search engine would tell people that and send them to your site. That’s the ideal, great for everyone. It lasted for a few years. Then came GoogleAds. Now not only would people get sent to your site, you’d also get paid for them being there. Even better, right? For a minute. But as soon as there was a financial reason for people to go to your site, someone else wanted that traffic instead. Now where your site was on the search results list became even more important and everyone was constantly optimizing to try and get on the first page of search results, and ideally the first result. Especially important if you were selling something on your site, not just getting paid by ads from people visiting your site. Then you could buy a “sponsored” result and be right at the top. OK, not great, but if it got people to your site and you made more from them going to your site than the ad cost then ok whatever.
But then other people started buying the sponsored results, and worse they put up websites filled with GoogleAds and just scraped all the content from your site and put it on theirs. Search engines didn’t care, content was content to them. After a bit of uproar Google and others made a policy to de-monetize sites that were just pulling content from others and slapping their own ads around it. For a moment it was better, then Google was like “You know what people hate? Having to click a link. Let’s improve user experience by offering summaries on search result pages!” Of course these summaries were just text pulled from your website, so now rather than some scammer trying to steal your traffic, the search engines started doing it themselves. This was bad, and then they added AI summaries and it got even worse. Entire categories of websites traffic dropped to zero because Google and others were showing people what was on the site rather than sending people to the site to see it. I don’t make my livelihood by writing on my blog, but if/when I ever look at my traffic it’s mostly crawlers and bots now, with the occasional person now and again.
This has largely been cat and mouse for decades. Update this file, change that setting, sitemaps, robots.txt, blah blah blah. One step forward, three steps back. In a new move being called “Independence Day” Cloudflare is launching new more powerful tools which will let site owners block these crawlers specifically the AI ones to try and claw back a little of the visit value websites hold. Importantly, and this is the big one, for the first time ever the new default also blocks Google search crawlers too. And that means the other search crawlers are on the block soon too. So if publishers have to consciously opt in to let search engines access their site, they simply aren’t going to. And that means how we find things online is about to be very different. Recommendations are trusted sources are about to get a lot more powerful.
Speaking of writing things online, Anthropic announced earlier this week that they are going to start putting watermarks into plaintext generated by Claude. This is to comply with some EU disclosure regulations, but as Gruber noted the announcement was light on details and brings up lots of questions. If you ask Claude to copyedit text you wrote, does it add these watermarks? If you quote text someone else used Claude to write in a larger piece you are writing yourself, is your piece going to get flagged for being written by AI? Anthropic clarified it a bit more a few days later, but this is already a touchy subject with any number of linguistic bits now being considered evidence of AI, even (and often) when AI isn’t being used. The em dash will never recover. There’s a bigger issue here which I’ve trying to write about for months (the irony of course is I could have AI write it all for me in minutes but I actually enjoy writing and working through my thinking on stuff which sometimes takes longer, so I do that) about how this policing around “use of AI” fully depends on lines that don’t exist, and results in just as many false positives as actual catches. You can debate if it’s good or bad, but I strongly suspect this panic will end up being be short lived. You can’t convince me that someone going to a recipe website to get instructions for cupcake isn’t going to bother worrying about if the 15 paragraphs of storytelling they aren’t going to read anyway was AI enhanced. (fun fact: No one goes to recipe sites for stories, but you can’t copyright recipes, so the recipe sites fill the pages with stories in an attempt tp make something proprietary)
Speaking of outrage, if you’ve been building a collection of movies on your Playstation it’s getting thinned out without your permission. Sony is pulling almost 600 titles over licensing issues, and there’s nothing you can do about it. Thought you bought them? Wrong, it was a license pretending to be ownership. This is not a new problem and we’ve seen it play out before but there was actually a time when buying something meant you owned it not that you were just allowed to use it for a little while, in the olden times we called that renting. This certainly plays into why “analog” is back in style, but there’s a reality to buying a DVD and knowing that you won’t wake up tomorrow to find it’s been removed from your library and the same can’t be said for “buying” a film from a streaming service. Actually owning the media you buy is powerful. People were conned by convenience into thinking because it’s digital it’s software and has to have a license but that’s a lie. We used to have MP3 collections to listen to digital music (some of us still do, though the discerning among us have moved to FLAC) but we let Apple and others convince us storing our stuff on their servers was better, and now it’s not our stuff anymore.
The funny thing is, there is actually a robust way to probably own digital items and that’s tokens on the blockchain. Yes, NFTs solve this smoothly. You can make fun of monkey picture bubbles, fairly, but all those people who bought monkey pictures still own them and don’t have to worry about a license expiring and their monkeys being deleted. If you can get past the hype and the drama and sit for minute with the tech, you realize the potential is vastly larger than most people image at the moment.
OK I’ve rambled enough and the kid just made garlic knots.
-s