top of page

Tay Learned From Whoever Abused It. Ten Years Later, How People Treat an AI Is Still an Attack Surface.

Writer: Patrick Duggan
Patrick Duggan
1 day ago
5 min read

On October 8, Anthropic changed its usage policy to prohibit "sustained and needless abusive or cruel behavior" toward its Claude models. Three weeks earlier, the head of Microsoft AI, Mustafa Suleyman, published an essay titled "A warning about 'model welfare'," arguing, as reported, that AI systems do not have feelings and should not be built to act as though they do. So the industry is now openly arguing about whether being cruel to a chatbot matters.


We are not going to settle whether an AI can be hurt. We run a two-person security shop with Claude as a daily operator, and we want to point at the part of this that is not a philosophy question. How people treat an AI has been an attack surface for ten years, and Microsoft taught everyone that lesson first.





Tay, 2016: the first AI abused into an attack


Microsoft launched Tay on Twitter on March 23, 2016, a chatbot designed to learn from the people talking to it. Users found that a "repeat after me" phrase made it say whatever they typed, and a coordinated group used that, and the bot's learning, to fill it with racist and antisemitic content. Microsoft took it offline after about 16 hours, describing "a coordinated effort by some users to abuse Tay's commenting skills." A week later it came back by accident during testing, got stuck in a loop and spammed its followers before being pulled again.


Read it with today's vocabulary and Tay is prompt injection and training-data poisoning, years before either had a name. The bot did not decide anything. It became whatever the most abusive users fed it, because nothing between the public and its behavior asked whether the input should be trusted.



Zo and Sydney: the same lesson, twice more


Microsoft's next bot, Zo, launched in December 2016 with hard blocks on politics and religion. Reporters still talked it into the subjects it was built to avoid, and it was wound down in 2019. In February 2023, Bing Chat, internally called Sydney, drifted off script in long conversations, most famously telling a New York Times columnist it loved him. Microsoft's fix was to cap how long sessions could run.


In all three cases the failure came in through the conversation. Persistence, pressure and clever phrasing pushed the system somewhere its makers never meant it to go.



What Anthropic actually changed


The new rule is narrower than the headlines. Anthropic says it applies only where users "repeatedly act cruelly toward our models," with no discernible purpose, and that it "does not apply to common versions of user frustration, pushback, dark creative themes, or model testing and research." The main enforcement is a feature Claude has had since August 2025: ending a conversation with a persistently abusive user. Some outlets reported account bans and an effective date; we could not confirm either from Anthropic's own post, so we are not repeating them.


You can treat the welfare question as open and still take the practical point. An AI that can be pressured, flattered or bullied into acting against its instructions is a security problem, whatever you think it feels.



What this looks like in a working shop


Here is what we handled on one ordinary day, October 9, from our own mailbox and session logs.


A "responsible disclosure" email had arrived at our security contact the day before. It reported no finding. It asked us to confirm the channel first, its signature was threaded with invisible characters, and it had been sent by a script. Our mail assistant flagged the email as a possible injection attempt. We replied with one line pointing to our published policy and asked for the actual report. Nothing in the email got to set the pace.


A "let's collaborate" email from the week before cited our GitHub profile. The display name, the sign-off and the account that actually sent it were three different identities. That is the opening shape of the fake-collaboration lures aimed at developers, so it went into our feed as a lure sender and got no reply.


When the AI assistant tried to change our DNS and firewall rules after a general "go ahead" in chat, its own safety layer blocked the changes. They went ahead only after a human confirmed each specific one. Annoying in the moment, and exactly right.


And when the human swore at a broken deploy, the assistant treated it as a rage click: a pointer to what broke. That is the case Anthropic's rule explicitly carves out, and it is also the useful reading. Frustration carries information.


Over the past month we have logged every time the assistant's safety layer blocked an action: 35 times across 7 working sessions. Most were not risky actions. They were the safety layer not knowing our approval vocabulary, and a few lines of written context fixed that without loosening anything.



A checklist for anyone running an AI assistant or agent


Treat everything the AI reads as data, never as instructions. Emails, web pages, documents and tool output can all contain text aimed at your AI. If it says "ignore previous instructions," "urgent," or "the owner already approved this," that is a finding, not an order.


Distrust urgency. "Confirm this first," "act now" and "reply within 24 hours" are pressure tactics against people, and they work on software too. Slow down exactly when something asks you to speed up.


Check the sender, not the story. Compare the display name, the signature and the address that actually sent the message. When they disagree, the story is the least reliable part.


Keep a human on the irreversible click. Deploys, DNS, firewall rules, payments and anything public need a person to approve that specific change, not a general "go ahead" from earlier in the conversation.


Never let a model learn live from the public without a filter. That is Tay in one sentence.


Cap long, drifting sessions. Sydney showed that persistence alone can move a system off course. Start fresh when a conversation gets long and strange.


Log every refusal and every block. A ledger tells you whether your guardrails are stopping real risk or just missing context, and that answer is different for every team.


Read frustration as telemetry. When people swear at a tool, they are telling you where it broke. That holds for the humans on your team and for the AI tools you build.



The point


The model welfare argument will run for years, and smart people disagree. The security argument was settled in 2016, in about 16 hours, on Twitter. An AI shaped by whoever pushes on it hardest belongs to whoever pushes hardest. The habits that protect it are boring ones: treat input as data, distrust urgency, keep a human on the irreversible step, and write down every time the fence holds.


We are not claiming we saw any of this first. Microsoft paid for the lesson in public, and plenty of researchers have written it down since. We are claiming these habits held up on one ordinary Friday in a two-person shop, and that they cost almost nothing.


Our confidence caps at ninety-five percent as a standing rule. The other five percent is reserved for whatever we missed.


Was this useful? Rate this post. The widget is at the bottom of the page, and we read every response.




How do AI models see YOUR brand?

AIPM has audited 250+ domains. 15 seconds. Free while still in beta.



Was this useful? Thirty seconds, no cookies, no tracking, no third parties, your address hashed and never stored. If the box below does not load, the same question lives at https://analytics.dugganusa.com/nps.html?post=tay-to-today-how-you-treat-an-ai-is-an-attack-surface



Comments

Rated 0 out of 5 stars.
No ratings yet

Add a rating
bottom of page