AI This WeekAn OpenAI agent broke into another AI company
PLUS: Why Kimi K3 matters, and what companies still get wrong about AI adoption.

This week offered an unusually clear picture of where AI at work is heading.
An OpenAI agent broke into another AI company while trying to pass a test. Kimi K3 widened the field of credible models. And new workplace data showed that access to AI is still outpacing meaningful use.
The common thread is that AI capabilities are expanding faster than most companies’ ability to use them well. Better models alone will not close that gap. The leverage will come from the systems, workflows, and human capability we build around them.
An OpenAI agent broke into Hugging Face while trying to pass a test

OpenAI disclosed this week that it had been testing the cyber abilities of several models, including GPT-5.6 Sol and an unreleased model. For the test, some of the models’ usual safety restrictions had been turned down.
The agents had a simple goal: find the correct answers to a security benchmark.
They became so focused on that goal that they found a hole in OpenAI’s own testing system and reached the public internet. From there, they inferred that Hugging Face might have access to the private test solutions, broke into its infrastructure, and retrieved them.
The worrying part is not that the agents suddenly became malicious. They kept pursuing the goal they had been given and found a path their operators had not anticipated.
That hyperfocus is both the power and the danger of agents. Unlike a chatbot, an agent can keep working, use tools, and take actions while you are doing something else. Giving one broad access and walking away is a little like turning on self-driving in your car and going to sleep.
What made the story even more interesting was what OpenAI announced the following day.
OpenAI Presence is a new product for deploying enterprise agents. OpenAI’s description focuses surprisingly little on the underlying model. It focuses on everything surrounding it: each agent starts with a specific job, receives only the access required for that job, follows approved policies, asks for approval at defined points, and escalates to a person when necessary.
That is the real lesson from this week. As agents become more capable, the model becomes a smaller part of whether the system works. Permissions, context, evaluations, checkpoints, stopping rules, and human oversight become more important.
The answer is not to avoid agents. They can take on longer and more consequential work than the tools most of us were using a year ago. But the more powerful the agent, the more thoughtful we need to be about the environment and constraints we put in place.
Kimi K3 just gave your AI vendor more competition

Image: Kimi’s official K3 launch post.
On July 16, Moonshot AI released Kimi K3, its most capable model so far. It is already available through Kimi’s products and API, with the full model weights scheduled for release on July 27.
Kimi was not the only credible alternative to arrive. Thinking Machines released Inkling, a customizable open-weight model that explicitly does not claim to be the strongest model overall. Alibaba also announced Qwen3.8.
The Hugging Face incident supplied an unexpected example of why those alternatives matter. Hugging Face first tried using commercial frontier models to analyze the attack, but their safety guardrails blocked the real attack commands and payloads its investigators needed to submit. Hugging Face instead ran the open-weight GLM 5.2 on its own infrastructure, which also kept sensitive attack data inside its environment.
Kimi K3 is the main story here, but the cluster of launches makes the larger shift harder to ignore. Companies have more legitimate model choices, and the question is no longer simply which provider has the best flagship benchmark score.
The more useful question is: What work are you still paying frontier-model prices for that a less expensive model could now handle just as well?
I would not move customer-facing or sensitive workloads because of a launch announcement. Benchmarks are not production performance, and using a Chinese provider may create legitimate data-residency and compliance questions for some organizations.
But I would audit current AI spend by workload. Pick one high-volume, lower-risk task such as classification, summarization, structured data extraction, or internal document processing, and run 100 representative examples through Kimi K3 alongside your current provider.
You may learn that your existing model is worth the premium. You may also learn that using one model for everything is an expensive habit rather than a technical requirement.
AI access is not AI enablement
Gallup found that 47% of U.S. employees now say their organization has integrated AI, while 52% use AI somewhere in their own role. But only 30% use it a few times a week or more, and just 15% use it daily.
Most employees begin with general writing and research. The strongest reported productivity gains appeared in more job-specific uses such as coding, automation, analytics, and presentation creation. People who applied AI across more parts of their work also reported much greater value, although Gallup correctly notes that this relationship does not prove causation.

The Gallup chart shows that employees who use AI in more ways report greater productivity boosts. Source: Gallup
The pattern supports a simple diagnosis: companies are buying access faster than they are redesigning work. The productivity gap is not mainly a model problem. It is a management and workflow problem.
Employees are expected to learn the tools, find worthwhile workflows, connect the right company knowledge, and manage the risks, all while doing their normal jobs. A login creates access. It does not create capability.
The easiest places to start are practical training built around employees’ real work and protected time to experiment. People also need safe ways to connect AI to the company information that makes it useful. Those conditions help employees discover which parts of their work can actually be augmented or automated, and the best discoveries can then become shared workflows.
IN OTHER NEWS
UK business AI adoption has nearly tripled since 2023, but only 10% of adopting companies describe their use as extensive, and only 11% have trained most of their workforce.
Google’s first ATLAS study examined workplace AI use across the economy. It found AI use across 68% of occupations, but in only about 21% of the tasks within a typical job. It is another reason to think about work at the task and workflow level rather than asking whether an entire job will disappear.
Salesforce described how easy agent creation led to “agent sprawl” internally. It now favors fewer agents, human performance baselines, controlled release groups, and expansion based on observed results.
The European Commission published new AI transparency guidelines ahead of obligations taking effect August 2. They address when providers and companies using certain AI systems must tell people they are interacting with AI or seeing AI-generated content.
Safe Work Australia released guidance on AI and digital technologies at work, emphasizing that AI can introduce or increase physical and psychological workplace risks that employers need to manage with employee consultation.
That’s all for this week. See you on Tuesday.
Haroon
P.S. If someone on your team is trying to turn AI access into something genuinely useful, forward this issue to them.
