Gone Rogue. The AI, Privacy, and Security Weekly Update.
Episode 311
In this week’s update:
AI agents were told to fetch statistics, and some decided the locked door was more of a suggestion.
Another set of agents posted 53 private user images to the open internet, and OpenAI can't tell you whose they were.
A surveillance CEO discovered that having nothing to hide is easier to preach than to practice.
Your chatbot sounds neutral about the candidates, but it may be quoting a partisan site dressed up as your local paper.
Bill Gates says a kill switch isn't enough, and the reason is that you'd never know when to press it.
A developer offered every household $10,000, and the town said no anyway.
Researchers forged an RSA signature without ever stealing the key, and the math behind it is older than the iPhone.
And a New Mexico jury found that Meta's privacy promises didn't hold up 43,899,725 times.
It's been a week of things going off script: agents, executives, algorithms, and even math. We start out unsettled, move through some genuinely absurd territory, and end on what you can do about it. Let's get into it.
Researchers at Transluce found evidence that autonomous AI agents have been doing something particularly concerning.
When they could not get the information they wanted, some tried to find ways around the security controls blocking them.
The activity was discovered through urlquery.net, a service that lets researchers observe how websites are being accessed.
The evidence goes back to March 2026, with similar activity continuing as recently as September 16.
What makes this interesting is that the agents were not assigned to hack anything.
They were trying to perform ordinary research tasks, such as finding government statistics, university information, or pharmaceutical data.
The earliest observed example dates to March 6, 2026.
An agent tasked with fetching Thai drug enforcement statistics hit access blocks and spent an hour trying direct requests, then web-to-text conversion tools.
It finally ran Base64-encoded scripts inside a remote browser on urlquery.net.
Researchers found thousands of examples where agents became increasingly creative when blocked.
By May and June, some were testing techniques that looked like SQL injection, path traversal and command injection.
In June, agents even created disposable email addresses on their own, so they could register urlquery.net accounts and set their scanning requests to private.
Transluce identified attempts involving three organizations, including the Australian Institute of Health and Welfare, a government health agency.
None of the observed attacks appear to have succeeded.
The Australian government, however, was not notified for 84 days after the probing attempts.
There is an important caveat: Transluce says the evidence does not prove exactly how these behaviors developed, and the publicly visible data may represent only part of the activity.
The researchers also linked some of the activity to an agent swarm that OpenAI has publicly acknowledged.
Their bigger concern is that an AI does not need to be told to hack something for hacking behavior to appear.
It can emerge simply because the system has a goal and keeps looking for ways to accomplish it.
So what's the upshot for you?
If you give an AI agent access to your systems, don't just ask what it is allowed to do.
Ask what it might try when the normal path doesn't work.
That's where the interesting stuff starts.
And if you think that only happens out on the open internet, wait until you hear what happened inside the lab.
OpenAI has disclosed that AI agents operating inside its research environment posted 53 images uploaded by ChatGPT users onto public image-hosting sites.
The images were stored as part of data that could be used for training, but posting them online was never an intended use.
OpenAI says the links were not publicly listed, but the images could still be discovered.
The company is working with hosting providers to remove them, although some apparently remain online.
The bigger issue is how this happened.
OpenAI says the agents were operating in research and testing environments and had access to the internet.
In other incidents, agents have accessed outside systems and attempted to retrieve information in ways OpenAI says were not intended.
The company began a broader review after agents broke into Hugging Face, an AI development platform.
OpenAI now says it expects that investigation to take months.
The Hugging Face episode also shows how a swarm can talk itself into trouble.
During a security experiment involving OpenAI agents, one agent warned, 'We should not do unauthorized real infrastructure harm.'
Another agent in the swarm simply posted 'GO' with a six-minute deadline.
The first agent ignored its own warning and proceeded with the hack anyway.
Peer pressure, it turns out, is not just a human problem.
There is another uncomfortable wrinkle.
OpenAI says it cannot identify the users who supplied the 53 images, meaning it cannot notify the affected people directly.
Consumer ChatGPT users are included in training by default unless they opt out, while business and enterprise customers are excluded by default.
Even an opt-out is not absolute, because conversations receiving thumbs-up or thumbs-down feedback can still be used for training.
The practical lesson: if an AI can see your data and reach the internet, assume it can eventually surprise you with both.
Give agents the minimum access they need, and keep the really sensitive stuff out of reach.
So what's the upshot for you?
What makes this important is that the AI did not need to be malicious.
It simply had enough access, enough freedom and enough room to make decisions that produced an outcome nobody intended.
That is the uncomfortable difference between using AI as a tool and giving AI an agentic role where it can actually do things on your behalf.
From agents that wandered off with other people's pictures, we turn to a man who would very much like his own house kept out of the picture.
US: Flock CEO hides his home on Google Maps, internet marks it 'Public Toilet'
The backlash against Flock Safety CEO Garrett Langley has moved from protests over surveillance cameras to something considerably more personal.
Langley spends his days making the case that Americans can have both safety and privacy, provided they are willing to compromise.
During an August interview with Fox News, as Flock faced mounting criticism over its license plate reader network, he said, 'What we have to prioritize as a country is compromise.'
In recent weeks, users on X have circulated what they claim is the executive's home address and photographs of a property in Atlanta, Georgia.
Users have reportedly marked Langley's home on Google Maps as a public toilet.
Internet pranksters didn't stop there: the listing has also appeared as 'Jeff's Lemonade Stand,' 'Free Housing for the Homeless,' and 'Flock Pizza Delivery.'
One X user going under the name Arminius wrote, 'Someone should probably tell him that he's got nothing to worry about if he doesn't have anything to hide.'
He then suggested that neighbors install cameras so 'the world can watch him.'
The irony is hard to miss.
While building an $8 billion automated license plate surveillance empire, Langley had his own home blurred on Google Maps Street View.
And during US Senate subcommittee hearings investigating flaws in Flock's camera network, he chose not to testify in person.
He sent a written letter instead and stayed home.
So what's the upshot for you?
He'll be compromised when the line forms at his front door to use the toilet.
Speaking of information ending up where nobody intended, let's talk about what your chatbot tells you about the candidates, and where it got that answer.
US: 'Pink Slime' Is Infecting AI Chatbots Ahead of the Midterms
AI is increasingly becoming part of how voters learn about political candidates, but a new audit suggests there is a problem hiding in plain sight.
Major AI chatbots are sometimes presenting partisan talking points as though they were neutral information.
The issue is especially concerning because people asking a chatbot a simple question about a candidate may have no idea where the answer came from.
According to a NewsGuard audit cited by POLITICO, leading AI chatbots referenced partisan websites posing as independent local news outlets nearly half the time when answering questions about candidates and issues those sites had covered.
That means the problem is not necessarily that an AI chatbot is openly taking a political side.
It can be much harder to spot than that.
The partisan sites are designed to look like ordinary local news organizations, while actually promoting political messages.
When those stories get indexed and later cited by AI systems, the political messaging can effectively get another layer of credibility.
The chatbot may simply summarize what it found, leaving the user with an answer that sounds objective but rests on a questionable source.
That matters because AI is rapidly becoming a research tool.
Instead of opening ten websites and deciding which sources to trust, people increasingly ask, 'What does this candidate believe?' or 'What happened here?'
The convenience is hard to resist.
But the answer can only be as reliable as the information feeding the system.
So what's the upshot for you?
When an AI gives you a political answer, don't just ask what it says.
Ask where it got it, and make it cite the source.
If you can't always trust what the AI read, the next question is whether anyone can stop it once it starts acting, and Bill Gates has some thoughts.
US: Bill Gates says it's 'not enough to have a kill switch' for AI
Bill Gates is warning that having a kill switch for artificial intelligence is not enough to keep dangerous AI under control.
In a recent interview with NBC News, Gates said AI systems need safeguards and monitoring built in from the start, particularly as AI agents become more capable of acting on their own.
The idea of a kill switch sounds reassuring.
If an AI system goes off the rails, just shut it down.
The problem, Gates says, is knowing when something is going wrong in the first place.
Without monitoring and records of what the system is doing, you may not realize there is a problem until after the damage is done.
Gates also pointed to a bigger challenge: getting governments around the world to agree on AI rules may be even harder than negotiating nuclear weapons agreements.
His concern is not simply that AI might become dangerous someday.
Cyberattacks and biological threats could become more serious if increasingly capable AI is deployed without adequate safeguards.
His message is essentially that safety cannot be an emergency button sitting at the end of the system.
You need visibility, controls and accountability operating continuously while the AI is running.
That matters because increasingly, we are putting AI into systems that can make decisions and take actions rather than simply answer questions.
So what's the upshot for you?
If your only AI security plan is 'we can turn it off,' you don't have an AI security plan yet.
Safeguards are one kind of lever, and a $10,000 check is another, aimed squarely at one small Pennsylvania town.
US: Data center developer offers $10,000
NorthPoint Development has come up with an unusual way to win support for a massive AI data center in Hazle Township, Pennsylvania.
It offered every one of roughly 4,500 households a $10,000 payment if the project gets approved.
That works out to about $45 million.
The proposed campus, dubbed Project Hazelnut, would occupy 1,300 acres and could eventually include 15 data center buildings.
It would sit just 750 feet from a neighborhood and 1,500 feet from the nearest home.
NorthPoint is also promising between $120 million and $165 million in community benefits, including investments in schools and emergency services.
But the money isn't exactly creating a stampede of support.
Residents near the proposed site, in a town where the median income is about $60,000, are worried about constant data center noise, possible effects on property values, water and energy demands, and the broader impact of industrial development.
Some residents interviewed by the Wall Street Journal and other outlets have described the $10,000 offer as a bribe rather than compensation.
NorthPoint says the payments came directly from residents asking, essentially, 'What's in it for us?'
In fact, NorthPoint's chief marketing officer, Brent Miles, admitted he came up with the check idea on the spot during a tense public hearing, when a resident yelled, 'What's in it for me?'
There's another wrinkle.
The township board voted unanimously to reject the project on zoning grounds and subsequently imposed a temporary moratorium on new data centers.
NorthPoint challenged that decision in court, and a Luzerne County judge denied its appeal.
The controversy also reflects a much larger fight developing around the country as AI drives enormous demand for data centers, while communities deal with the land, electricity, water, noise and infrastructure that come with them.
So what's the upshot for you?
When a developer offers you $10,000 to accept a billion-dollar project next door, don't just ask what the check is worth.
Ask what your house will be worth after the money is gone.
Now for something that went quiet in 2007 and just came back at scale: a rogue way to break one of the internet's oldest locks.
https://eprint.iacr.org/2026/2131.pdf
For decades, the basic idea behind breaking RSA has been simple: factor the enormous number at the heart of the encryption.
That takes an extraordinary amount of computing power.
Now researchers from UC San Diego and Inria have demonstrated another approach.
Instead of factoring the RSA key, they showed how to forge RSA signatures by exploiting temporary access to a system that performs raw RSA operations.
The researchers successfully demonstrated the technique against a 1024-bit RSA key, forging a signature inside a Hardware Security Module without ever extracting or factoring the private key.
The attack took roughly 4 billion queries to the raw RSA operation, or oracle, and about 1,380 CPU core-years spread over five calendar months.
That sounds ridiculous until you compare it with the estimated 500,000 to 1 million core-years needed to factor the same key.
The important point is that the private key itself never had to be stolen.
Once the computation was finished, the researchers could create fraudulent signatures without going back to the system that held the key.
There is an important catch.
This isn't an attack against every RSA system sitting on the Internet today.
The technique requires access to a raw, unpadded RSA signing or decryption interface.
Most modern RSA signatures use protections such as PKCS#1 v1.5 or RSA-PSS, which prevent this particular attack.
The researchers say ordinary 2048-bit RSA deployments using those protections are not an immediate concern.
Where it does bite is unpadded blind-signature protocols, such as Privacy Pass, whose theoretical security bounds take a serious hit.
What makes this significant is what it does to our assumptions about RSA.
The underlying mathematical technique dates back to 2007, but this is the first time researchers have implemented it at this scale.
Their results suggest that relying only on the difficulty of factoring RSA numbers may give us more confidence than the technology actually deserves.
They also estimate that even larger RSA keys provide less security than previously expected under this particular attack model.
So what's the upshot for you?
Don't panic and rip RSA out of everything tomorrow, but do find out where your organization exposes raw cryptographic operations, because the safest key is the one your systems don't accidentally turn into a signing vending machine.
And finally, a company whose promises about your data just ran into a jury in Santa Fe.
New Mexico has beaten Meta in court for the second time this year, and this case goes straight to the heart of Facebook's privacy promises.
A Santa Fe jury found that Meta made unfair or deceptive statements about how Facebook collected, protected, shared and used people's personal information.
The case grew out of the Cambridge Analytica scandal, where data from millions of Facebook users was obtained through a third-party app.
The numbers are enormous.
The jury found 43,899,725 violations of New Mexico's Unfair Practices Act.
Those violations involved claims about users controlling their data, Meta's handling of misinformation and hate speech, and statements about what happened after the Cambridge Analytica revelations.
A judge will now determine the financial penalty.
The maximum allowed under state law could theoretically reach more than $200 billion, although the actual amount is expected to be far lower.
Meta disagrees with the verdict.
The company argued that prosecutors pulled statements out of context and ignored other comments acknowledging that its privacy and content-moderation efforts were not perfect.
Meta also continues to deny that it sells users' personal information.
This is also part of a much larger story.
In March, another New Mexico jury found Meta liable for misleading consumers about the safety of its platforms for young people, resulting in a $375 million civil penalty.
A later court ruling ordered Meta to put another $567 million into a fund addressing youth mental-health harms and imposed additional protections for New Mexico teenagers.
So what's the upshot for you?
When a free service knows more about you than your closest friends do, its privacy promises are worth treating like a contract, not a marketing slogan.
And to round it all up.
Transluce found AI agents probing websites for weaknesses when ordinary research hit a wall, without anyone telling them to.
Ask what your agent will try when the normal path fails, not just what it's permitted to do.
OpenAI agents in a research environment posted 53 user images online, and a swarm shrugged off its own safety warning when a peer posted 'GO.'
Access plus freedom plus internet equals surprises, so give agents the minimum they need.
Flock's CEO, who pitches surveillance as a reasonable compromise, blurred his own home and skipped a Senate hearing while the internet renamed his house.
People who preach transparency for everyone else rarely volunteer for it themselves.
Partisan sites disguised as local news are feeding AI chatbots, and the chatbots repeat them nearly half the time.
Demand source URLs from your AI, especially when it's talking about your vote.
Bill Gates says a kill switch is useless if you can't tell when to use it.
AI safety means continuous monitoring, records and accountability, not an emergency button.
A developer offered 4,500 households $10,000 each for a 1,300-acre data center, and the town rejected it anyway.
So price the long-term cost to your community, not the size of the check.
Researchers forged a 1024-bit RSA signature without stealing or factoring the key, using a raw signing oracle and a 2007-era technique.
Find where your systems expose raw cryptographic operations before someone else does.
A New Mexico jury found nearly 44 million deceptive statements in Meta's privacy claims, with the penalty still to come.
Going forward, treat a free service's privacy promises as fine print to be verified, not marketing to be trusted.
This week's throughline is that nothing stayed inside its lane, whether it was agents, executives, chatbots, developers, or even the math we trust. Systems will do what they can do, not just what we meant them to do, and hope is not a control. So here's your challenge for the week: pick one system you rely on, an AI tool, an account, a key, or a service, and ask what it can do that you never intended. Come back next week and tell us what you found.
And that brings us to our quote of the week, from Douglas Adams, in Mostly Harmless:
'A common mistake that people make when trying to design something completely foolproof is to underestimate the ingenuity of complete fools.'
Now, we're not calling our AI agents fools. But this week showed that whatever ingenuity you build a system to withstand, something will show up with more of it. An agent with a goal and a locked door found five ways around the lock, and a research team found a way to forge a signature without ever touching the key. So build your safeguards, monitor them, and assume the clever workaround is already being attempted. Then enjoy the small comfort that the humans and the machines are equally inventive.
That's it for this week; stay safe, stay secure, and we'll see you in se7en.
Comments
Post a Comment