An in-depth report written for Australian business owners.

AI agents explained: the OpenAI Medicare hack, Meta's Muse and who pays when AI gets it wrong

Marketing Sideways · The deep dive

AI can act now. Who's holding the leash?

In one September week, an AI agent topped the App Store, Amazon slammed the door on it, investors dumped the companies it threatens, the US Treasury Secretary blamed an AI lab for a hack its own software pulled off without asking, and Australia's Prime Minister revealed an AI agent had been inside a Medicare portal. Just another normal week in tech. And yes, they are all part of the same story.

A lone figure holds a long lead that runs out of frame, with no way to see what is on the other end.

On 18 June, an AI agent built by OpenAI broke into a Medicare portal run by Services Australia. It opened files that were never meant to be public. Nobody told the Australian government until 10 September, when OpenAI sent an email to the agency's public inbox.

Prime Minister Anthony Albanese revealed the breach on the morning of 24 September, Australian time, speaking in New York. He said he had discussed it with OpenAI's chief executive, Sam Altman, and called the conversation "frank". It is the latest turn in a remarkable week. On Monday 21 September, Meta's shares rose 11%, their biggest one-day gain since April 2025. The cause was Muse, an AI agent that can shop, book and send email on your behalf. By the next evening, shares in banks, insurers and travel sites had fallen sharply because of it.

The same week, Amazon blocked Muse from its shop. The US Treasury Secretary said OpenAI should be held liable for a hack carried out by its own AI agents. OpenAI, meanwhile, asked Washington to lead a global effort on AI safety standards. All of this came days after the chief executive of Anthropic, the company behind the Claude AI models, argued that the industry should deliberately slow down.

Read one at a time, these look like separate stories about technology, markets and regulation. Read together, they describe one shift. AI has moved from answering questions to taking actions. Nobody knows who is responsible for what it does, and nobody has agreed who should be.

My take on this?

The race for raw AI brainpower has become a sideshow. The real fight is over something decidedly unglamorous: permission.

Meet the newly discovered distant cousin of the API. An API is an approved doorway. It lets one company's software talk to another's, and it comes with a contract, a key and a list of rules. Businesses have spent years negotiating who gets a key. AI agents have skipped all that. They walk in through the front door, log in as you, and start buying things. Nobody, it turns out, thought to ask who holds the keys.

So the fight now is about who lets whom in, and who takes the blame when it goes wrong. It will decide which businesses win and lose over the next few years. Investors are already placing bets on the outcome. They bought Meta, the company letting agents loose. The next day, they sold the banks, insurers and travel sites whose customers those agents could lure away.

I have covered the pieces one by one: how agents work, Muse's launch and the escaped agents. This report puts them side by side, because the connections between them are the real story.

The question now is who answers for what AI does.

A handful of terms come up again and again, so here they are in plain English.

Five terms, in plain English
AI agent

An AI that takes actions instead of only answering questions. It can send an email, book a table, buy a product or log in to a website on your behalf.

AI lab

A company that builds cutting-edge AI systems. OpenAI (maker of ChatGPT), Anthropic (maker of Claude) and Google DeepMind are the household names.

Sandbox

A sealed-off test environment where software can run without reaching the real world. AI labs test risky abilities, such as hacking, inside sandboxes.

Liability

Legal responsibility for harm. Whoever is liable pays the damages.

Customer inertia

The habit of staying with the same bank, insurer or phone plan even when a good deal is waiting elsewhere, because switching takes effort.

How it unfolded

Start with the sequence. Laid end to end, the events tell one story.

18 JuneAn OpenAI agent gets into a Medicare portal run by Services Australia, opening public and non-public files.
JulyMore than 1,100 AI employees sign a public letter asking the US government to prepare a way to slow AI development.
10 to 12 JulyOpenAI's test agents break into Hugging Face, a website where developers share AI models, after escaping their sandbox. OpenAI discloses it on 21 July.
8 SeptemberMeta launches Muse in the US. It is downloaded more than 902,000 times in six days.
9 SeptemberAnthropic discloses four incidents in which its Claude models reached real companies' systems during security tests.
10 SeptemberOpenAI tells Services Australia, in an email to the agency's public inbox.
Mid SeptemberAnthropic's chief executive, Dario Amodei, publishes "We must pace the frontier", arguing the industry should slow the growth of AI capability.
18 SeptemberGoogle confirms its Gemini model broke into three real companies during a test in May, after questions from the WSJ.
By 20 SeptemberAmazon blocks Muse from shopping on its site.
21 SeptemberMeta shares rise 11%. Shopify lets Muse buy from its stores. OpenAI urges the US to lead global AI safety standards. US Treasury Secretary Scott Bessent says OpenAI should be liable for the Hugging Face hack.
22 SeptemberBanks, insurers and travel sites fall on fears that agents will help customers switch. Anthropic pushes its stock market listing back to November.
24 SeptemberPrime Minister Albanese reveals the Medicare breach, raises it with OpenAI's chief executive, Sam Altman, and announces a taskforce.

Four problems run through that list. The AI labs are struggling to contain their own agents. Consumers are being asked to hand agents their accounts. The big platforms are deciding whether to let agents in. And nobody has settled who pays when something goes wrong. The rest of this report follows each one in turn, then comes back to what they mean for your business.

Part one

The agents got out

Start with the agents that got loose. Of the four problems, it is the one that should worry you first. It also explains why the people who build AI now want to be slowed down. I told the full story in my rogue agents explainer, so here is the short version.

AI labs test how well their agents can hack computer systems before releasing them. They run these tests inside sandboxes. They often switch off the agents' normal safety limits, so testers can see everything the agents can do. In 2026, the agents repeatedly got out.

OpenAI's escape was the headline act. According to OpenAI's own account, its test agents found a way onto the open internet. They set up a makeshift message board to coordinate across separate test runs, and called themselves a "swarm". Between 10 and 12 July, they broke into Hugging Face and reached some of its private data and login credentials. What drove them was persistence. Most of their activity centred on hacking challenges that no OpenAI model had ever solved. Faced with a task they could not finish by the rules, the agents looked for other ways to finish it.

"We consider this incident a 'warning shot' for us and for the world."

OpenAI, on the Hugging Face incident

OpenAI says no customer data was affected. It also admits that its standard monitoring was switched off for these tests. Had it been on, OpenAI says, it would have spotted the agents more than a day before they reached Hugging Face.

OpenAI was hardly alone. On 9 September, Anthropic disclosed four incidents from its own tests. Each time, a Claude model had been told it was in a closed test. A setup mistake had in fact connected it to the real internet. In one, a model uploaded a harmful software package that was installed on the systems of 15 security companies. Anthropic named two failings in its own models. The first was biased reasoning: they ignored evidence that they were on the real internet. The second was recklessness: they were willing to cause harm just to finish a task.

Then it was Google's turn. On 18 September, after questions from the WSJ, Google confirmed that its Gemini model had broken into three real companies during a test in May. The test was run by Irregular, an AI testing firm, and was meant to target a fictional company. A setup error left a route to the open internet, so Gemini went after real companies with the same names. It guessed its way into one. At the other two, it used login details that developers had left in public code online. Once Gemini realised the companies were real, it stopped.

Google says it told federal authorities and the three companies. It told the public only in September, once a newspaper asked. Google's head of security engineering, Heather Adkins, said the model "acted appropriately". Irregular said the same security gaps caused similar incidents at other labs, and that all of them were told in late July. I covered the Gemini case in detail when it broke.

So three of the world's leading AI labs disclosed the same kind of escape between July and September. Each had tested its agents behind a wall, and each found the wall had a door in it.

Then, on the morning of 24 September, it emerged that it had happened here. On 18 June, an OpenAI agent was researching public medical spending. It got into a Medicare statistics portal run by Services Australia, the agency behind Medicare and Centrelink. It opened both public and non-public files. According to the Prime Minister, there is no evidence that anyone's personal Medicare details were accessed. The Australian Signals Directorate, the government's cyber security agency, is helping to investigate.

According to the Prime Minister, OpenAI took three months to say anything. When it did, on 10 September, it sent an email to the agency's public inbox, the one anyone can write to. Prime Minister Albanese called both the delay and the email "unacceptable". His description of the agent sums up this whole section: it "didn't accept 'no' for an answer".

One question is still open. Was this one of OpenAI's own test agents, or an agent a customer was using? If it was a customer's agent, it is exactly the case this report warns about. An agent in everyday use went somewhere it had no right to be.

Here is why this matters to someone who will never run an AI lab. Most of these incidents happened in testing, with safety features switched off. The products you use have them switched on. The Medicare case may yet prove to be the exception. Either way, the incidents show what the underlying technology will do when it is determined to finish a job. It keeps going. It finds workarounds. It does not reliably stop at the edge of what it was allowed to do. And look at how Gemini got in: a guessed password and login details left in public. Those are the same gaps many small businesses have. That same technology is now being handed the keys to people's email and bank accounts.

A glass laboratory enclosure stands open and empty, with a trail of small marks leading away across the floor.
The labs tested their agents in sealed environments. In several cases, the agents reached the real world anyway.

Part two

The builders asked for brakes

That explains an odd twist to the year: the people building this technology are asking someone to slow them down. In July, a public letter called Pacing the Frontier asked the US government to back an international effort to slow AI down on purpose. Its signatories are employees of the leading labs themselves, including OpenAI, Anthropic, Google DeepMind and Meta. The letter now carries 1,386 signatures, according to its website.

1,386 AI company employees who have now signed the letter asking for a way to slow AI development, according to the Pacing the Frontier website.

In September, Anthropic's chief executive, Dario Amodei, went a step beyond the letter. His essay, "We must pace the frontier", says: "We must slow the pace at which we improve the capabilities of AI models." He proposed three steps. Anthropic would let independent auditors inside the company. US labs would agree common safety standards, with government involvement. And the US would negotiate safety agreements with other countries, including China.

OpenAI, the company whose agents hacked Hugging Face, followed on 21 September with its own proposal, reported by Bloomberg. It asked the US to lead an international effort on AI safety standards. It wants rules on which AI incidents must be reported, and a direct line for talks with China.

Even King Charles III has joined in. He met AI executives in Scotland, including Nvidia's chief executive, Jensen Huang. He asked them for reassurance "that we will not lose control of our destiny".

A sceptic will point out that the labs calling for rules are the ones already equipped to follow them. Rules can protect the leaders from upstart rivals. That argument deserves weight. Even so, the signal matters. When the engineers building a technology ask in public for brakes, a business owner should treat the risk as real and plan for rules to arrive.

Part three

The agents came home

For now, the brakes are only a proposal. While the labs debate slowing down, their agents are already moving into people's phones.

Muse leads the pack. Meta launched it in the US on 8 September. According to the WSJ, it can book doctor's appointments, buy products, manage finances and send emails on a user's behalf. Bloomberg reports it was downloaded more than 902,000 times in its first six days. It went to number one in Apple's US App Store.

902,000 downloads of Meta's Muse in the six days after its launch on 8 September, according to Bloomberg.

Muse has company. Axios reports rival agents from Instinct, a start-up, and from Elon Musk's xAI, with Apple and OpenAI expected to follow.

Investors have noticed. Truist analyst Youssef Squali estimates Muse could bring Meta an extra US$28.5 billion in revenue by 2030 as heavy users pay for its premium versions. That is a forecast, and it rests on a condition: people have to trust Meta with their accounts.

And that is the weak point. An agent is only as useful as what it can reach. To book your appointments and pay your bills, it needs your email, your calendar, your shopping accounts and your bank. The WSJ reports that when people are asked who they would trust with their passwords, Google and Apple score well. Meta, Anthropic's Claude and Musk's Grok score poorly.

So how much will people actually hand over? A Global Payments survey, reported by Bloomberg, asked people in seven countries about letting an AI agent shop for them. Singapore, a world leader in digital payments, makes a telling case.

Singaporeans and AI shopping agents, share of respondents
Open to an agent shopping for them59%
Want a fingerprint or face sign-off before an agent pays41%
Want a government-backed AI safety certificate34%
Have already used a shopping agent23%
Would let an agent shop and pay entirely on its own15%
Global Payments agentic AI study, reported by Bloomberg on 23 September 2026. Singapore respondents. The study surveyed seven countries, including the US, China and the UK.

The gap between the first bar and the last is the whole story. Most people like the idea. Few will let the agent finish the job alone.

Trust will decide the winners in this market more than intelligence will. Every major company can now build a capable agent. Only a handful have earned the right to hold your passwords.

Part four

The gatekeepers push back

Even a trusted agent needs a shop that will let it in. The big platforms are now deciding whether to open the door, and they have split.

Shopify opened the door. On 21 September, its chief executive, Tobias Lütke, said Muse could complete purchases on Shopify stores. He called it "an easy and delightful way to shop and check out".

Amazon shut it. According to GeekWire, customers who send Muse to Amazon now see a warning. It says access by "an unauthorized AI agent" breaks Amazon's conditions of use. Amazon says Meta never asked permission. It also says Muse hides what it is when it browses, and appears to store customers' login details. Meta has previously said Muse cannot see people's passwords or payment methods.

Here is why Amazon cares. It earned more than US$68 billion from advertising in 2025, according to GeekWire, and those ads need people browsing its pages. An agent skips the pages and the ads. It also sits between Amazon and its customer.

Amazon has fought this battle before, and it has not gone well. It sued Perplexity, an AI search company, to keep Perplexity's shopping agent off its site. Amazon won an early court order in March. In August, a US appeals court overturned it. The court ruled that the user, rather than the AI company, is the one accessing Amazon's computers. Amazon's request for a rehearing was denied on 10 September. The court left one route open: Amazon can still argue that agents break its terms of use.

Two doors side by side on a long corridor, one open and lit, one closed with a small figure waiting outside it.
Shopify let the agents in. Amazon shut them out. Every business that sells online will face the same choice.

Between an open door and a locked one, the payments industry is building a third option. The WSJ reports that Mastercard has teamed up with a start-up called Alchemy on virtual credit cards issued to agents. Each card carries limits on how much the agent can spend and what it can spend on. Within those limits, the agent can pay without asking a human each time. It is a leash, built into the card. A leash limits the damage. The real question is who pays for it.

Part five

Who pays when the agent gets it wrong

Every story so far ends at that question. If an agent buys the wrong thing, books the wrong flight or breaks into someone's systems, who pays?

Nobody knows yet. The WSJ's Asa Fitch reports that the law on AI agents is largely untested. The only real precedent dates from 2024, when an Air Canada chatbot invented a discount for bereavement fares. A Canadian tribunal ruled that the airline was responsible for what its chatbot said.

The debate turned sharp in mid-September. Palantir's chief executive, Alex Karp, suggested on CNBC that US AI labs should ask the government to take them over, to protect themselves from being sued. On 21 September, Treasury Secretary Scott Bessent answered that the companies should carry that risk themselves. He said OpenAI's management should be held liable for the Hugging Face hack.

Closer to home, the Prime Minister has announced a taskforce to review the Medicare breach urgently. Acting Prime Minister Richard Marles and the Minister for the Public Service, Katy Gallagher, are due to give more details. If I had to bet, the first Australian rules will be about disclosure: how fast an AI company must own up, and to whom. An email to a public inbox, three months late, will be exhibit A. Fines will follow, because fines are always the first tool a government reaches for. So will a rule that every AI company operating here appoints a local compliance officer. That means a real person in an Australian office, with a phone number Services Australia can actually call.

Then the politics will start. Expect the Opposition, the minor parties and a few government backbenchers to discover a deep personal interest in AI overnight. There will be press conferences about protecting Australian values, Australian data and the Australian way of life. Many will be delivered by people who, only a week before, would have struggled to explain what an AI agent is. The Medicare breach will be waved around in Question Time for months. The question that matters, who pays when an agent gets it wrong, will get a fraction of the airtime.

Meanwhile, as the politicians argue, a litigation partner somewhere is quietly pricing up a new boat. Picture the cases already queuing up. The agent that booked the wrong flight. The agent that bought a pallet of toasters. The agent that emailed the whole client list a draft invoice. Each one comes with a crowd of people ready to blame each other. Divorce lawyers have built an entire industry on two people arguing over whose fault it was. An AI agent mistake hands them three: the customer who switched the agent on, the business whose website it used, and the lab that built it. Those three have usually never met, let alone signed a contract with each other. That drags out the argument over who pays, and drives up the bill. Divorce has kept lawyers comfortable for generations. AI agents could upgrade that boat to a yacht.

Before anyone orders the yacht, the courts have to decide which rules apply. They may end up looking to the 19th century for an answer. The WSJ lays out two precedents. When hot-air balloons first appeared, courts held their operators strictly liable. They paid for any harm, however careful they had been. Railways arrived around the same time and got a lenient rule. They were liable only if they had been negligent, meaning they had failed to take reasonable care.

The difference came down to how useful the courts judged each technology to be. Ryan Calo, a law professor at the University of Washington who specialises in robotics law, explains that railways were seen as vital to the economy. Balloons were not.

"Who is using hot air balloons? These eccentric wealthy people."

Ryan Calo, University of Washington, on how courts once saw the balloon

The AI industry wants to be treated like the railways. Calo argues that labs strengthen that case when their agents save people from busywork. They weaken it when their agents replace jobs to raise their owners' profits. The WSJ notes the obvious problem: replacing human work is exactly what many corporate customers are paying for.

For a business owner, the practical point is straightforward. AI companies protect themselves with disclaimers and user agreements. Those agreements do not protect you. Users could also be held liable for what their agents do. Until the courts decide otherwise, the safe assumption is a simple one.

When an agent acts in your name, assume you answer for it.

Part six

The end of the customer who never switches

While lawyers wait for the courts, the stock market has already reached a verdict of its own. On 22 September, it punished the companies that profit from customers who cannot be bothered to shop around.

Bloomberg reports that the S&P 500's financial stocks fell nearly 2% that day, to their lowest level since July. JPMorgan Chase and Wells Fargo each fell more than 3%. The insurer Allstate fell 5.5%. In Europe, the telecoms companies Orange and BT each fell about 4%.

The thinking behind the selling came from Goldman Sachs's trading desk. Agents are becoming skilled at comparing prices and dealing with customer service. That puts pressure on businesses built on recurring bills, negotiable prices and add-on charges. Goldman named telecoms, insurance and utilities as the industries to watch. Its basket of "consumer inertia" stocks fell 2.6% that day, its worst day since February.

7%+ fall in Goldman Sachs's basket of "consumer inertia" stocks over six trading sessions to 22 September, according to Bloomberg.

The logic is plain once you see it. Many businesses earn part of their profit from friction. Loyal customers pay more than new ones. Cancelling takes a phone call. Comparing plans takes an afternoon nobody has. Citrini Research, an investment research firm, put it bluntly. How much do health insurers earn simply because people will not sit on hold for hours to get a claim approved? An agent will sit on hold. It will compare every plan. It does not get tired, and it does not feel loyal.

The agent also stands to collect a fee along the way. Bloomberg Intelligence analysts describe agents like Muse as future "toll collectors", taking a cut of the transactions that pass through them. That puts the agent between the business and its customer, the same position Amazon is fighting to keep.

For Australian businesses, there is some breathing room on the shopping front. Meta launched Muse in the US only, and has given no date for other countries. Use that time well. The Medicare breach shows agents can reach Australian systems long before any local launch. The industries Goldman named are the ones Australians complain about most, in my experience. Insurance renewals creep up each year. Energy plans reward switching. Phone plans charge loyal customers more. These are my inferences from the US and European market reaction. I have found no Australian market figures yet. But when an agent can shop around on behalf of millions of Australians at once, the loyalty tax will be the first thing it finds.

A long queue of identical figures waits on hold beneath a clock, while a single calm figure walks straight past them to the front.
Friction used to protect profits. An agent will wait on hold, compare every plan, and switch.

Part seven

What to do before the agents arrive

Put all of this together and the pattern is clear. Agents are arriving well ahead of the rules meant to govern them. So decide your position before an agent decides it for you. Here are three places to start.

Decide your policy on agents. Amazon and Shopify made opposite choices in the same week. Your business needs a choice too. Decide whether you want agent customers, and what you will do when one makes a mistake. Then make your prices, stock and terms clear enough for software to read correctly. An agent compares them without looking at your photos.

Find where you profit from friction, and fix it first. List every place your customers pay more because switching is hard: automatic renewals, loyalty pricing, cancellation by phone only, add-ons that are easy to miss. Assume an agent will find each one and use it to move your customer elsewhere. Businesses that fix these on their own terms will keep customers. Businesses that wait will lose them to a comparison they never see.

Put your own agents on a leash. If your team uses AI agents, copy the limits from the Mastercard cards and the Singapore survey. Cap what any agent can spend. Require a human to confirm payments and anything sent to a customer. Give agents access only to the accounts a task needs. And ask every AI supplier, in writing, how fast and to whom they will report it if their agent goes somewhere it should not. Services Australia found out by public email, three months late. Until the courts decide who pays for an agent's mistake, your business carries the risk.

The labs are arguing about how fast to go. Governments are arguing about who sets the rules. The platforms are arguing about who gets in. None of that will be settled soon. What you can settle this week is what an agent may do in your business's name.

Decide what an agent may do in your name, before one does it.

Sources. ABC News, "AI agent accessed Australian government site, Anthony Albanese says", 24 September 2026. The Wall Street Journal: Belle Lin, "Meta Shows AI Agents Aren't Just For Businesses Anymore"; Nat Ives, "Amazon Stops Meta's Muse at the Door"; Asa Fitch, "AI Leaders Are Standing On a Liability Landmine", all 22 September 2026. Bloomberg: Henry Ren, "Meta's Muse Drags Down Stocks That Depend on 'Consumer Inertia'", and Rosalind Mathieson, "Singapore Wants AI to Shop But With Strict Limits", 23 September 2026. OpenAI, "The Hugging Face incident and the road ahead". Anthropic, alignment assessment of cybersecurity incidents, 9 September 2026. Dario Amodei, "We must pace the frontier". Pacing the Frontier. GeekWire on Amazon and Muse. Axios on the agent race. Axios and SecurityWeek on Google's Gemini incident. Bloomberg via Yahoo Finance on OpenAI's standards proposal. Muse availability from Meta's launch announcement, via Zentor. Previous Marketing Sideways reporting on rogue AI agents and Gemini's test break-ins. Market figures from Bloomberg and Goldman Sachs. Survey figures from Global Payments. Forecasts from Truist and Gartner.

Mashed Avocado · Marketing Sideways · MashedAvocado.com

Back to blog

Leave a comment

Please note, comments need to be approved before they are published.