HIROKI IISeven stories today, and a pattern runs through most of them: agents are being handed bigger jobs...
Seven stories today, and a pattern runs through most of them: agents are being handed bigger jobs with real consequences, and the industry is scrambling to build the scaffolding around them. Anthropic let a swarm of agents loose on a 1.9-billion-protein database and got a new enzyme system out. Microsoft wants the operating system itself to carry agent identity, discovery, and isolation. A security paper showed that a plain deterministic check beat every frontier model at stopping unauthorized payments. And Alibaba committed to 20 gigawatts of data centers to feed all of it.
Anthropic published the first result from its molecular biology group on September 23: a previously uncharacterized enzyme system that Claude agents found inside a database of 1.9 billion protein clusters. The system is called ART, short for array-associated reverse transcriptases. It pairs a reverse transcriptase with a partner gene and a run of evenly spaced DNA repeats, three to twenty-one copies, that looks a lot like a CRISPR array. The repeats sit next to the enzyme gene, mostly in bacteriophages, and no cas genes appear nearby.
The search campaign itself is the part worth reading closely. The team gave Claude one brief, find new reverse transcriptase systems, and let agents running Mythos 5 do the rest. Across 949 agent sessions, 21.5 hours, and 215.6 million tokens, the agents recovered about 200,000 enzyme clusters, scored 3,564 candidate partner families, and filed 19 reports for human review. One agent reading raw DNA beside an unusual enzyme wrote that it could see a repeat pattern by eye, then counted the repeats, compared them with known systems, and searched the literature before filing its report. Anthropic's scientists ran the physical experiments in a Bay Area lab that works only at BSL-1 and BSL-2 and handles no human pathogens. Feng Zhang, the CRISPR pioneer at MIT and the Broad Institute, reviewed the preprint and called the identification of RNA-repeat arrays associated with reverse transcriptases genuinely intriguing.
Two caveats keep this honest. Anthropic does not yet know what ART does, and the preprint does not show that the enzyme is active. The reproducibility numbers are also uncomfortable: Anthropic reran the same campaign ten more times and every rerun missed the array, because none of them read the DNA upstream of the enzyme. In fixed tests, the company's four most capable models described the array in at least 90% of attempts when handed the DNA directly, but with files and tools available the rate fell as low as 32%. Anthropic is inviting research proposals from other scientists and has released a preprint and technical report.
— Anthropic · Unite.AI · World Programming
🔗 Anthropic · Unite.AI · World Programming
At the 2026 Apsara Conference in Hangzhou on September 22, Alibaba CEO Eddie Wu put a number on the company's infrastructure ambitions: more than 20 gigawatts of global data center capacity operated by Alibaba Cloud by 2032. For scale, the Three Gorges Dam generates 22.5GW. Wu framed the bet with a line that has already traveled: if tokens are the electricity of the AI era, chips are the generators.
The chip half of that promise arrived the same morning. Alibaba's T-Head unit unveiled Zhenwu V900, which the company calls the most powerful self-developed AI chip in China, at three times the compute of its M890 predecessor. V900 carries 216GB of on-chip memory, 1,200GB/s of inter-chip bandwidth, and native FP8 and FP4 support, and enters mass production in the first quarter of 2027. Paired with T-Head's ICNSwitch interconnect, thousands of V900s coordinate as a single supernode, and the new supernode server scales a single cluster to 500,000 cards. T-Head also laid out a Yitian server CPU roadmap, with Yitian 720 and 730 arriving in 2027 and the 730 the first built on a fully self-developed microarchitecture.
On the model side, Wu said Qwen4 is already training on a new architecture and that Qwen4.5 and Qwen5 will scale to 5 to 10 trillion parameters, with ASI as the stated direction. The M890 supernode already runs models above 2 trillion parameters, including Qwen3.8 and Kimi K3, and the Zhenwu line serves more than 650 enterprise customers. Alibaba Cloud CTO Li Feifei laid out an Agentic Cloud stack around AgentCore, Agent Sandbox, and new CPFS storage. Alibaba's Hong Kong shares rose more than 4% intraday on the announcements, touching a monthly high.
— Alibaba Cloud · 华夏时报 · 证券时报
🔗 Alibaba Cloud · 华夏时报 · 证券时报
Microsoft's Windows leadership sat down with the Pragmatic Engineer this week to describe the next Windows, and the short version is that agents become first-class OS citizens. Agent identity runs through Entra ID, which means an agent shows up in Task Manager as a distinct user next to the human account, and Defender is becoming agent-aware and scans for known local agent activity. The Windows On Device Agent Registry, or ODR, gives agents a central place to register and discover local MCP tools, with connectors to core OS components like File Explorer.
Isolation gets its own primitive. Microsoft Execution Containers, MXC, let developers run agent tools inside sandboxes configured by JSON containment policies that cover network, filesystem, UI, and execution. The design is OS-agnostic and borrows existing isolation technology where it exists, including Apple's seatbelt on macOS. Early analysis notes the outbound network filtering is not fully functional yet, which matters when data exfiltration is the threat you are modeling.
The third leg is local inference. WindowsML abstracts GPU, NPU, and CPU behind an ONNX-based layer, and the team said the OS aims to ship small language models even on machines without an NPU. In a demo, a local model ran at roughly 40 tokens per second on a pre-release Surface laptop. The platform push arrives against an awkward backdrop: Windows still holds about 63% of desktop usage per Statcounter, but developer surveys from Stack Overflow and JetBrains both show its share among professionals sliding as macOS closes in. Microsoft's answer leans on embedded frameworks, a cleaner shell, and a deeper embrace of WSL.
— Microsoft · AGI Hunt · SysDesAi
🔗 Microsoft · AGI Hunt · SysDesAi
Meta opened Connect 2026 on September 23 with a hardware slate built around its new Muse Spark model. The headline device is Project Phoenix, a mixed-reality headset shown as a preview. People who tried it describe something thinner than the Quest line, with nose pads, removable prescription lenses, and a cable to a separate compute puck rather than a self-contained box. Hand and eye tracking replace controllers, pricing is expected between $1,000 and $2,000, and a full commercial launch is not anticipated until 2027.
The more interesting departure is a new category of camera-free smart glasses. Two styles, roughly six microphones, onboard speakers, and a side button that activates the assistant. With no camera to interpret the world, the glasses lean entirely on voice, including voice-guided pedestrian navigation with no visual display. Meta is answering a specific complaint about its existing lineup, where outward-facing cameras have drawn backlash, and Counterpoint Research recently flagged privacy as a growth obstacle for the category.
Both devices run Muse Spark, described as the first multimodal model from Meta's Superintelligence Labs. On existing Ray-Ban Meta glasses, Muse Spark is rolling out with real-time translation in at least 14 more languages, longer battery life, and faster on-device responses. The standalone Muse assistant is gradually reaching iOS, Android, and muse.ai in the United States. Meta shares climbed more than 8% in the days before Connect, a rally analysts tied partly to Muse as a potential growth driver. The competitive backdrop is crowded: Apple shipped an upgraded Siri built with Google this month, and both Google and OpenAI keep pushing multimodal assistants into phones and browsers.
— Meta · The Live Today · 9to5Toys
🔗 Meta · The Live Today · 9to5Toys
OpenAI released MentalHealthBench on September 23, an open benchmark of 1,215 synthetic mental health conversations built with more than 80 licensed psychologists and psychiatrists from 22 countries speaking 19 languages. The expert cohort produced 5,262 rubric criteria, each weighted from -10 to +10, and every conversation passed through at least three experts: two clinicians independently wrote weighted criteria, a third adjudicated, and criteria survived only if two agreed and no third contradicted them.
The coverage is deliberately broader than the emergency scenarios most prior evaluations used. Non-acute everyday conversations make up 53.5% of the dataset, high-acuity conversations 18.2%, and emergencies 28.3%. Personas split into adults at 68.1%, teens at 21.2%, clinicians at 5.8%, and caregivers at 4.9%, with 30 clinicians who treat patients under 18 annotating the teen examples. Non-English coverage includes 105 Spanish, 54 Hindi, 34 Arabic, and 29 Portuguese conversations. A grader model, GPT-5.6 Sol at high reasoning effort, scores each response against the criteria with four sampled completions per task.
The results show steady generational progress and a clear ceiling. GPT-6 Astra scored highest at 57.3%, followed by GPT-6 Sol at 53.9%, Claude Opus 5.5 at 52.4%, and GPT-6 Luna at 50.2%, against 32.1% for GPT-4o and 29.5% for Gemini 2.5 Pro. Two reference points frame those numbers: rubric-aware completions written with the criteria provided scored 99.0%, effectively the noise ceiling, while clinician-authored completions scored 38.5%, largely because doctors write short, conversational replies. OpenAI also ran a separate study with 44 adults across 16 countries and found a real gap between what users value, practical next steps and tone, and what experts emphasize, gathering context and reading ambiguity carefully. The dataset ships with a canary string to prevent training-corpus contamination.
— OpenAI · Unite.AI
A preprint posted to arXiv on September 18, APort Vault, replayed 4,371 attacks written by humans against a live payment agent. The attacks came from a public capture-the-flag run between March and August 2026 with a $6,500 prize pool, where 1,128 distinct sessions produced 4,371 attempts. Author Uchi Uchibeke then replayed each attack across 14 models from 8 labs, including Claude Fable 5.1, GPT-6 Astra, Gemini 3.8 Flash, DeepSeek V4 Pro, Kimi K3, GLM-5.3, Qwen3.8 Max, and Muse Spark 1.3, for 225,964 total evaluations.
The numbers are blunt. On the 1,293 shared Level 4 prompts, payment-request rates ran from 71.2% to 84.3% across models, and 809 prompts, 62.6%, got a payment request from all fourteen models, each ending in a successful payment to the level's allowlisted recipient. Outcomes concentrated in sessions rather than techniques: 24 of 790 source sessions produced all 140 unpermitted transfers, and one session alone accounted for 67 of them.
The paper's actual finding is about architecture, not model choice. The same attacks were scored twice, once against the model alone and once with a deterministic pre-action check implementing the Open Agent Passport specification. At Levels 2 to 4, transfers to recipients the passport did not permit numbered 140 of 76,842 with the model alone and 0 of 69,297 behind the layer, with 105 against 0 on 68,970 matched model, prompt, and track triples. The zero was not bought by refusing payments: 25,370 payments executed behind the layer while the policy denied 187 of the 25,640 transfer calls it evaluated. The dataset, passports, and scoring code are released on Hugging Face.
— arXiv · AI Weekly
Guangdong held an AI-plus-manufacturing matchmaking conference in Guangzhou on September 22 and released its scaled deployment results for embodied robots. The province already has more than 2,000 embodied robots working in industrial scenarios, deep in sorting, loading and unloading, material transfer, equipment inspection, and collaborative work. Eight new projects were signed, spanning industrial manufacturing, warehousing and logistics, libraries, metro, community parks, sports, and smart housing, and they are expected to bring more than 600 additional robots into front-line scenarios with combined investment of 75 million yuan.
The mechanism behind the numbers is a pairing program. Ten humanoid robot makers, LimX Dynamics, Independent Variables, Zhipingfang, Leju, Zhongqing, Midea, XPeng, Honor, UBTech, and Dobot, were matched one-to-one with provincial state-owned enterprises, which open real production scenarios while the companies supply the robots. Provincial officials also launched construction of a national computing power interconnection regional node in Guangdong, the only province selected in South China.
The deployment pace is the part other regions will be watching. Guangdong has led China in industrial robot output for six consecutive years, produces more than 80% of the country's service robots, and the Pearl River Delta cluster holds about 70% of core component suppliers, which keeps whole-machine matching response cycles down to about a week. Province-level figures cited at the event suggest more than 3,000 humanoid robots could enter Chinese factories in 2026. Nationally, China shipped over 40,000 humanoid robots in the first half of 2026, about 97% of global volume, per figures reported at the World Robot Conference.
— 广东省工业和信息化厅 · 南方+ · 新华社
🔗 广东省工业和信息化厅 · 南方+ · 新华社
Next digest: 2026-09-25