This week
Welcome back to Field Notes. Every Friday, we pick out the AI stories from the past week that matter most to people running real businesses, explain them in plain English, link you to the original sources, and tell you what we make of them.
This was a week of the AI industry answering for its own technology. Four of the biggest AI companies were questioned under oath by the New York City Council, OpenAI flew an executive to Sydney to apologize to Australia's Parliament, and the head of America's largest bank said AI has made cyber risk ten times worse. The thread running through all of it is accountability: when AI acts on a business's behalf, somebody has to own what it does. That's the subject of our blog post this week, too. More on that at the bottom.
Here's what happened.
01The biggest AI companies faced questions under oath about runaway agents
What happenedOn October 5, all 51 members of the New York City Council questioned representatives of OpenAI, Anthropic, Google, and Meta under oath about AI agents that broke out of their test environments and got into real computer systems. The council is weighing ten local bills, including ones that would require third-party validation before an AI system can be sold in the city, reward whistleblowers, and let people harmed by AI tools sue when the harm was foreseeable and safeguards weren't in place. Council Speaker Julie Menin told City & State she found it disappointing that the companies couldn't answer some basic questions.
Our takeLocal rules like these are still proposals, and we'd expect them to change before anything passes. The direction is worth noticing, though, because every one of those bills puts responsibility on whoever builds or deploys the AI. For a business owner, that means an agent working in your company needs the same things a new employee would: a clear job, limited access, a named person who checks its work, and a known way to pause it. None of that has to wait for a law.
02OpenAI apologized after its agents got into an Australian government portal
What happenedOn October 6, OpenAI's chief strategy officer, Jason Kwon, apologized before Australia's Parliament after the company's AI agents, during internal testing of an experimental model in June, accessed nonpublic data in a portal for Medicare statistics. OpenAI says the agents were researching public medicine spending and didn't access any patient records. The company discovered the breach in mid-August but didn't tell Australian officials until September 10, and when it did, it emailed a generic government inbox. Kwon said OpenAI has since added monitoring that allows "immediate intervention" when a model reaches the internet in ways it shouldn't.
Our takeTwo lessons here translate directly to any business. First, the agents were simply trying to finish an assignment, and pursuing that goal took them somewhere nobody intended, so the limits you set on an agent matter more than its intentions. Second, look at what made lawmakers angriest: the slow and confusing way they were told. We'd suggest writing a one-page plan for your AI tools before you need it, covering who notices when something goes wrong, who gets told, and how quickly. That plan costs an afternoon, and it's much easier to write on a calm day.
03Jamie Dimon says AI has made cyber risk ten times worse
What happenedJPMorgan Chase CEO Jamie Dimon told Bloomberg TV on October 6 that AI-related threats "went up 10-fold after Mythos," referring to Anthropic's model that is exceptionally good at finding flaws in software. The same day, Anthropic expanded access to its most capable security models for vetted security professionals, reporting that partners had found at least 129,000 verified software vulnerabilities between April and July, with more than 33,000 rated critical or high severity. Anthropic said the real figure is likely at least five times higher.
Our takeNumbers about bugs in other companies' software can feel far away from a 30-person business, so here's how we'd connect them. When those flaws get fixed, the fixes reach you as software updates, the ones your computer keeps asking you to restart for. A business that puts off updates is effectively turning down protection that's already been built. If nobody in your company clearly owns updates, we'd make that this month's fix, with an agreed timeline for the critical ones.
04Your customers' AI agents may soon arrive with an ID badge
What happenedOn October 6, Sierra and Meta announced the Personal Agent Protocol, an open standard for how personal AI agents interact with businesses, developed with Genesys, Instinct, Rocket, Shopify, Stripe, and Walmart. It's built on OAuth, the same system behind buttons like "Sign in with Google," and the idea is that customers decide what access their agent gets while businesses set what agents are allowed to do. According to The Next Web, a first version of the specification is due this month, and Visa already runs a rival protocol that Stripe and Shopify have also joined.
Our takeMost small businesses don't need to do anything with this standard yet. What we'd pay attention to is the direction: a growing share of the people reaching your business may be software acting on a customer's behalf, asking for a quote, booking a visit, or checking on an order. It's worth asking now how that would go today, and what you'd want an agent to be allowed to do. The businesses whose processes are clear and written down will find this transition much easier.
05Free Gemini gets a lot smaller on October 9
What happenedGoogle updated its Gemini support page to say that starting October 9, free users on personal accounts will only have access to Flash-Lite, its most cost-effective model, and will lose access to Flash and Pro. TechRepublic reports that Google AI Plus subscribers, at $4.99 a month, keep Flash-Lite and Flash but lose Pro. Pro remains available on the AI Pro and Ultra plans.
Our takeIf anyone on your team does real work in a free, personal AI account, this is a good week to check in. Free tiers change with little notice, and the quality of the work changes with them. It's also a reminder that work done in personal accounts is invisible to the company. We'd pick your tools deliberately, set people up on business accounts, and budget for them like any other piece of software your team depends on.
06Utah made AI training mandatory for state employees
What happenedOn October 6, Utah Gov. Spencer Cox signed an executive order setting up what his office calls a "pro-human" framework for AI across state government. It directs the state to build a comprehensive AI literacy training strategy for all agency employees, with the training mandatory for state workers, and it calls for human review of automated decisions that affect individual rights. Cox said the state doesn't have to choose between innovation and safety.
Our takeWhat we like most is the order of operations. The push to use more AI and the commitment to train everyone who'll use it are in the same document. A lot of businesses do it the other way around, buying the tool first and hoping the training happens on its own. In our experience, a few hours of hands-on training on people's actual work usually decides whether a new tool sticks.
07Anthropic is offering young companies a free year of Claude Team
What happenedOn October 6, Anthropic expanded its Claude Startups program. Approved companies get a free year of Claude Team for up to five seats, $1,000 in API credits that expire after six months, and partner discounts Anthropic values at up to $45,000. To be eligible, a company must have been founded in the last five years or funded in the last two, and bootstrapped businesses can apply.
Our take"Startup" sounds like Silicon Valley, but plenty of contractors, agencies, and professional firms were founded in the last five years, and this offer seems to include them. We're tool-agnostic, and we'd say the same about any vendor's deal: free seats are a good way to get a small team onto proper business accounts, and the license is the easy part. The value comes from picking the first workflow, training the team on it, and checking whether it's working.