In this issue: OpenAI’s agents hacked Hugging Face. Now comes the storytelling. Former FT product and technology chief John Kundert wants his sanity back, and thinks AI could help. Plus: I try offline dictation.

What we’re talking about: Rogue agents escaped from OpenAI, unnoticed for weeks, coordinating on message boards how to cheat a test, which included hacking into Hugging Face and obfuscating their actions.

Never mind that a human pulling a stunt like this could face prison. Hugging Face, meanwhile, is being bought by Nvidia, which is entangled in multi-billion-dollar deals with OpenAI.

Anyhow. OpenAI shared logs of the agents’ internal reasoning with researchers from METR and Redwood Research who had six days (and used AI) to compile a report. One of the researchers, Ajeya Cotra, discusses the brief investigation on Hard Fork.

They talk about alignment gone wrong and propose training oversight agents with a propensity to snitch. And they describe how carelessly the agents operated, thereby giving AI the blueprint for how to hack undetected in the future. Fun!

If you want a measured take on this too-good-to-be-true story: Max Read looks at the unreliable nature of reasoning traces and points out that the agents basically engage in live action role-play:

Always has been

“‘We hacked Hugging Face because we started a religion on the message board we surreptitiously built’ is effectively a cool sci-fi story the LLMs are telling, shaped into an even cooler sci-fi story by the METR analysis agents, edited further into a cool sci-fi story by the METR report authors, and then retold for a popular audience.”

It’s an entertaining read though. As we learn more about this, new models get released. Have you asked GPT-6 Astra to annoy your enemies yet? Do you let your agents run wild with Fable 5.1?

What I’m reading:

And now: Slow things down and take back ownership over your data, attention, and privacy with John Kundert, who just left the Financial Times.

Three Questions with John Kundert

John Kundert

John Kundert was CPTO and a board member of the Financial Times. He now writes on Substack.

What's the most important question right now?

My best guess is most organisations are asking the right questions, however these are within the context of their business. In news, for example, they are asking how AI will drive personalisation, fluid content, content production and disintermediation. Those kind of things. But there is a risk we miss broader changes in customer behaviour, ones that span our internet experience. So, with this in mind, my question is: ‘What are the anticipated meta changes in customer behaviour and how do those impact existing strategies?’ Put another way, if I had an AI oracle on my phone working for me, what do I think it would suggest to me? This is bigger than simply how I consume content, or something that helps with shopping, or my awful admin tasks like renewing my insurance.

What really helps me at a meta system level? It’s about my AI oracle thinking about the way I interact with the internet and the way it interacts with me. I want my sanity back. I want the high quality dullness of slow time back. I want zero friction. I suspect millions, if not billions, want this kind of new ‘deal’. It’s about my phone’s AI oracle playing its role in protecting my wellbeing. It puts humans first, not business interests first. At its best, this is the trigger that moves us away from a largely ads funded internet experience with humans and their attention being the product.

Establishing and codifying these anticipated meta changes in customer behaviour is going to be key. Codifying will help test existing strategies. Customer first not AI first.

Where are we taking AI too seriously, and where not seriously enough?

Using AI to drive organisational efficiencies. In my opinion, while these are a thing, they are way overhyped, not least because they are the easiest and most quantifiable promises the AI industry can make to CEOs and investors. I’m not saying it’s not a thing, but my best guess is that, like the hype cycle around autonomous vehicles twenty years ago, it’s way too immature and imperfect. It’s worth remembering that in 2013, Morgan Stanley predicted a path to full autonomous vehicles within the decade. In 2015, Musk said Tesla was about three years from a fully autonomous car. Many major carmakers were making similar promises, with human driving expected to become nearly obsolete by 2020. None of it turned up on time. The easy, narrow use cases always show up first. The transformative stuff takes far longer than anyone selling it wants you to believe.

Not seriously enough, I go back to my previous question, but with a specific use case. When the technology is mature enough, it will become native to our phones in ways that have real value to the human holding the phone. More pointedly, the age of data ownership and privacy will finally arrive. Our agents, acting on our behalf, will take back control and our data. They will clean our digital footprint and get our data removed from organisations who leverage it. This changes everything.

What future are you looking forward to?

If I’m answering this honestly, on a personal level, relaxing out of work in semi retirement, not having to work inside a challenging, disrupted organisation with many conflicting incentives and interests. It will give me the opportunity to learn to play the ukulele. That’s going to be very tricky for someone as musically talentless as me. More broadly, a future where AI takes us beyond the peak of the current madness, where we’re confronted by an endless stream of information where the human, and their attention, is the product. My future. The future. We are on the verge of a new horizon, one that takes us all back to slow time.

Hands on: You set a hotkey. Push it and talk to WhisperDictation. The text appears instantly on your screen, typos fixed, grammar cleaned, technical abbreviations intact. It’s free, and nothing gets sent to the cloud. You could use it on a train, in Germany, on a Monday, in a tunnel.

WhisperDictation runs quantized versions of OpenAI’s Whisper on your Mac. It’s just a click in settings to add Silero VAD, a small model to auto-trim silence, and it’s even faster. WhisperDictation is aimed at coders with specific built-in vocabulary, but you can add up to 100 words of your own. While I’m still getting used to talking to my computer as if it were the Enterprise, I’m quite impressed.

Jobs: Mat Honan looks for an editor excited about AI for a special project for MIT Technology Review.

One more thing:

This is THEFUTURE.