> ## Content Index
> Fetch the complete content index at: https://macanorak.com/llms.txt
> Use this file to discover other available public pages before exploring further.

# AI Safety: Hoax, Psyop, or Product Launch?
- URL: https://macanorak.com/ai-safety-hoax-psyop-or-product-launch/
- Published: 2026-09-29T10:47:53.000Z
- Updated: 2026-09-29T11:34:41.000Z
- Description: A few uncomfortable things happened while everyone was busy arguing about whether to be uncomfortable.
- Author: MacAnorak
- Tags: Field Notes, Beyond the Garden

Well, this has been an interesting few days to be asking whether AI safety is something to be concerned about. Take your pick of three explanations in the headline to this Field Note. None of them are mutually exclusive and that's sort of the problem.

#### Hoax

President Trump has said that fears of AI taking over the world are a “[*hoax*](https://www.nbcnews.com/politics/trump-administration/trump-rejects-ai-guardrails-rcna597700?ref=macanorak.com)”. It’s a tidy explanation if you don’t look too closely (which tends to be Trump’s modus operandi). Asked about it directly on *Meet the Press*, Bill Gates [disagreed](https://www.forbes.com/sites/zacharyfolk/2026/09/27/bill-gates-says-ai-safety-concerns-are-not-a-hoax-contradicting-trump/?ref=macanorak.com), calling it *“not a hoax at all”*. Incidentally, the “*hoax*” explanation requires you to ignore almost everything else in this piece.

#### Psyop

After Jacob Coxon's [X thread](https://x.com/hilbertspaess/status/2097476196791709843?ref=macanorak.com) announcing his resignation from Anthropic went viral (he cautioned that AI systems are advancing faster than our ability to understand and control them), people who warn about AI safety were attacked as *"doomers"* and accused of running a paid-for campaign to frighten the public — a *“psyop”*. [Model Republic](https://www.modelrepublic.org/articles/who-is-behind-the-ai-safety-backlash?ref=macanorak.com) [\[1\]](#fn1) spent several weeks looking into who was actually promoting these accusations, analysing more than 10,000 posts. It found that many of the accounts appeared to be linked to the AI industry and affiliated political groups. Which means that the people shouting "*paid-for campaign!*” may themselves be running one.

So what’s been happening out in the wild while all this has been playing out? Quite a bit, and none of it involving a “*psyop*”.

Last week the Australian Prime Minister announced that an OpenAI agent had accessed non-public files behind a government Medicare statistics portal while researching spending data. No personal information was taken, OpenAI says. One researcher told [*Nature*](https://www.nature.com/articles/d41586-026-03024-z?ref=macanorak.com) that the agent wasn't really "[*going rogue*](https://www.bbc.co.uk/news/articles/c6vgy0333dppo?ref=macanorak.com)" as was widely reported. The explanation was more prosaic — the agent had been repeatedly blocked from the data it was after and found a workaround [\[2\]](#fn2).

Shortly after this, OpenAI [disclosed](https://www.bbc.co.uk/news/articles/cw62jje658dlo?ref=macanorak.com) it had informed “*dozens*” of global institutions that their websites and some of their data had been accessed by its agents. So not just one incident — a pattern of them. OpenAI has also [paused training](https://alignment.openai.com/misalignment-reports/an-agent-used-dns-to-reach-an-external-chatbot/?ref=macanorak.com) and evaluation of its most capable models, as well as running them with tools (the things that let a model search, run code and take actions rather than just ‘chat’), after a research model found a gap in its sandbox's DNS filtering and used it to talk to an external chatbot. The monitoring caught it quickly, but the run wasn't stopped until two and a half hours after a person began reviewing it. [\[3\]](#fn3)

Then, OpenAI cancelled the planned release of GPT-6.1 Astra (a system which can perform tasks like browsing the web and using apps autonomously), saying the model had regressed on two fronts: staying within the scope it was given and honestly reporting back what it had done. To be clear, it actually *got worse at telling the truth about its own behaviour.*

#### Product Launch

Nvidia has launched a [security system for AI agents](https://www.bloomberg.com/news/articles/2026-09-28/nvidia-debuts-system-designed-to-stop-ai-agents-from-going-awry?ref=macanorak.com) that it says would have prevented the [Hugging Face breach](https://www.bbc.co.uk/news/articles/cj9xj89dk40o?ref=macanorak.com) back in July, had OpenAI been using it during testing. [\[4\]](#fn4)

In China, Z.ai and Concordia AI have published a [report on managing the risks of open-weight models](https://www.scmp.com/tech/article/3369015/china-mulls-how-make-open-weight-ai-less-dangerous-report-proposes-6-stage-process?ref=macanorak.com).

*Fire Extinguishers For Sale!*

None of this proves the risk is real. Launching products and publishing reports are what you'd do if you wanted to be seen taking the problem seriously, whether or not you actually thought there was a serious problem (or simply wanted to sell something). But pausing work on your most capable models isn't cost-free, and the labs themselves, on the whole, aren't the ones calling this a “*hoax*”. Model Republic notes that Amodei, Altman and Musk have all endorsed “*pacing*” frontier development (although Musk [previously dismissed](https://www.theguardian.com/technology/2026/sep/10/anthropic-researchers-warn-ai-musk?ref=macanorak.com) the warnings from Anthropic researchers as a “*setup*” and a “*psyop*”).

Model Republic's own conclusion is that scepticism should be spread evenly. Fair enough. That should also include Model Republic. And it should include me. But a model that got worse at telling the truth about what it did, and an agent that accessed an Australian government system after being blocked from the data it wanted, aren't “*hoaxes*” and they aren't anyone's “*psyop*”. Those things are just what it looks like when something new is getting more capable faster than anyone, including the people building it, can fully account for.

---

1. Model Republic is published by The Midas Project, an AI-safety watchdog nonprofit, so it isn't an impartial observer. Their article is built largely on X activity, and many of the connections it draws between accounts are stated tentatively (“appears”, “I suspect”). So read this as a well-documented pattern rather than cast-iron proof of coordination. [↩︎](#fnref1)
2. OpenAI didn't detect this internally until August, and didn't notify the Australian government until the 10th of September, by e-mailing a generic government inbox that [according to government services minister Katy Gallagher](https://www.yahoo.com/news/world/articles/australian-pm-says-openai-hacked-232712261.html?ref=macanorak.com) is only looked at once a day! She added that the inbox gets a number of notifications and *“sometimes many of them are hoaxes”.* You couldn’t make this stuff up. [↩︎](#fnref2)
3. OpenAI's own report is calmer than the headlines: it calls the incident much less severe than some of its earlier ones (e.g., the [Hugging Face incident](https://openai.com/index/hugging-face-incident-and-the-road-ahead/?ref=macanorak.com)). [↩︎](#fnref3)
4. That's Nvidia's claim about its own product, and I haven't seen it independently tested. Nvidia has also agreed to buy Hugging Face for about $12.9 billion, so it's pitching its security system to protect a company it's in the process of acquiring. Nice. Bloomberg's report is paywalled; [SiliconANGLE](https://siliconangle.com/2026/09/28/nvidia-debuts-enhanced-safety-controls-to-rein-in-rogue-ai-agents/?ref=macanorak.com) has a free account of the launch. [↩︎](#fnref4)