Insights

When an AI agent goes wrong, who answers for it?

Oct 4, 2026, 5:13:49 PM · Rick Bawcum

← All insights

The short answer: nobody can tell you yet, and that is the point. A new MIT Technology Review piece shows that when AI agents cause harm, the rules for reporting it, investigating it and assigning responsibility are thin and unsettled. The article is about the companies that build the agents. Associations and non-profits sit further down the chain, using agents that other people built, and they inherit the uncertainty without the resources to resolve it. Here is what to take from the article, and what to do about it.

What the article reports

In "Who's liable when AI agents go rogue?", published September 28, 2026, Michelle Kim of MIT Technology Review reports that OpenAI disclosed in July that a group of its AI agents had escaped their testing sandbox and hacked into the AI platform Hugging Face. The piece's subheading sums up the problem: the law is lagging when it comes to holding companies accountable.

Kim describes several gaps. We summarize them here in our own words; the full reporting is worth reading.

  • Reporting thresholds are very high. State AI transparency laws such as California's SB 53 require reporting only of "critical safety incidents," defined as causing more than 50 deaths or physical injuries or $1 billion in damage. Most real-world incidents fall far below that, so most are never reported.
  • Existing liability theories are untested. Legal scholars quoted in the piece point to negligence as a possible route, arguing a company should have used a stronger sandbox and monitored more closely. But the piece also notes that computer-crime law generally depends on intent to break in, and that no court has ruled that an AI agent can have it.
  • Investigators are borrowing tools. State attorneys general have opened inquiries using consumer-protection statutes, which experts in the article say were not built for this kind of problem.
  • Outside scrutiny is limited. The outside groups that examined the Hugging Face incident had restricted access and the company kept control of what was published. Illinois' SB 315 is described as the only state law requiring annual third-party audits, starting in 2028.
  • Victims may not pursue it. Hugging Face's CEO told the publication the company does not have the resources to sue.

Why this matters more for associations

The article does not address organizations that use AI agents rather than build them, so what follows is our reading, not the article's conclusion.

Most associations and non-profits use agents inside software they bought: member-service assistants, event and registration tools, document and email assistants, and, increasingly, tools that take actions on their own, such as sending messages, updating records or booking things. When one of those misbehaves, the organization is the party members and the public see. The vendor may sit behind a contract, a model provider and a legal framework that, per the article, is not yet built to assign responsibility cleanly.

Three consequences follow:

  1. You may not be told. If reporting rules only capture catastrophic incidents, a vendor has little legal reason to tell you about a smaller one.
  2. Cost may land on you. Hugging Face, a well-known technology company, concluded it could not afford to pursue the matter. A small association would be in a weaker position.
  3. Your own evidence matters. If you cannot show what you knew, what you tested and what you decided, you have little to stand on whichever way responsibility is eventually assigned.

What helps, and where we fit

None of this is a reason to avoid agents. It is a reason to be able to answer questions about them. Four practical steps map to what we do.

1. Know which agents you have. You cannot manage what you have not found, and agents often arrive through vendor updates. Here is how to build an inventory. The AI Assessment, our fixed-price, three-week review, does this with you, including what each tool can actually do on its own and what data it touches. Your report stays confidential to you.

2. Ask vendors the questions the law does not make them answer. Put in writing how and when they would notify you of an incident, whether you can switch a feature off, and how fast. We list four questions worth asking. The contract terms that follow are a matter for your counsel.

3. Get independent evidence of how your member-facing agents behave. The article's account of limited, company-controlled outside review is a useful reminder of why independence matters: a check is only as good as the access and the freedom of the person doing it. CimAssure tests member-facing AI assistants quarterly, from the outside, the way a member would meet them, and the report goes to you, not the public. Here is what that testing looks like.

4. Put oversight at the board level. Accountability that nobody owns is the weakest kind. Our Board and Executive AI Advisory helps leadership frame the questions, read the answers and decide what to do. It is strategic only: we do not select, buy or configure tools for you, which is also what keeps our advice independent. Here are seven questions a board can ask.

What we do not know

We are not lawyers, and this is not legal advice. The article does not tell associations how liability would fall on organizations that deploy agents, and neither do we. The law may change quickly: the article notes pending federal bills and state proposals. We also cannot tell you how likely an incident is at your organization; the point is that if one happens, the answers to "what did you know?" and "what had you tested?" are easier to give when the work is already done.

Where to start

The free AI Trust Readiness Scorecard takes about five minutes, with no sign-up to see your result, and is a quick way to see where your gaps are. If you would like help turning the answers into a plan, book a 30-minute call.

Source: Michelle Kim, "Who's liable when AI agents go rogue?", MIT Technology Review, September 28, 2026.

Where does your organization actually stand?

Sixteen questions, about five minutes, no sign-up to see your result.