• NATURAL 20
  • Posts
  • AI Agents Get Faster and the Stakes Get Bigger

AI Agents Get Faster and the Stakes Get Bigger

PLUS: Google expands cyber AI, Figure locks in $3.5 billion of compute, and Moonshot files for a Hong Kong IPO.

Sponsored by

Smarter CRM. Less Busywork.

Disconnected data and tools make it harder to understand your customers. HubSpot's Agentic Customer Platform brings your data, teams, and tech stack together with AI built in to help your business work faster and create more personalized customer experiences.

Why HubSpot and what's new

  • Use AI powered tools to take action faster

  • Unify your data, teams, and tech stack in one place

  • Create one shared view of customer data

  • Connect teams around the same customer context

  • Bring your business tools into one place

Connect more of your business in one place and give every team a smarter way to work. Get set up quickly and start checking off your hardest tasks.

Today:

  • GPT-6 Astra cuts computer-task time by 47%

  • Muse Spark 1.3 uses fewer tools and tokens on coding work

  • Gemini 3.8 Flash pairs agent upgrades with a cyber model

  • Thinking Machines Lab: Nvidia considers roughly $2.5 billion investment

  • Sanders and Casar: Proposal would ban artificial superintelligence

AI Agents Get Faster—and the Stakes Get Bigger

OpenAI, Meta, and Google are pushing AI agents toward longer and more consequential work across computers, software development, and cybersecurity.

At the same time, the money behind those systems is getting larger. Thinking Machines is discussing a multibillion-dollar financing, while Figure has committed billions to the computing infrastructure needed to train its humanoid-robot models.

The result is an AI race increasingly defined by three things at once: capability, control, and compute.

Bar chart titled “ARC-AGI-3” comparing model scores: GPT-6 Astra leads with 99.9%, followed by Claude Opus 5 at 30.2% and GPT-5.6 Sol at 7.8%.

OpenAI released GPT-6 Astra for computer use, browsing, software engineering, science, cybersecurity, and professional work. It is initially rolling out to a limited set of organizations, with ChatGPT Plus, Pro, Business, and Enterprise access, the API, Microsoft Azure, and AWS Bedrock following over the coming days.

Astra can work across browsers and desktop software to fill forms, update customer records, research information, create documents, analyze data, install and test software, and check websites it builds. In OpenAI’s OSWorld 2.0 simulation, Astra scored 72.6% at roughly 40 minutes per task, compared with GPT-5.6 Sol’s 65.7% at roughly 75 minutes—about 47% less time in OpenAI’s test.

The model also scored 97.6% on FrontierMath Tier 4, 99.9% on ARC-AGI-3, and 100% on ExploitBench in OpenAI’s evaluations. These are OpenAI-run or OpenAI-reported results and may differ from production performance.

Cybersecurity is the most sensitive part of the release. OpenAI classifies Astra at the Critical level under its Preparedness Framework and reports that it discovered and used two previously unknown vulnerabilities during testing. The public version refuses advanced requests such as producing proof-of-concept exploits, while less restrictive defensive access is planned through OpenAI Daybreak.

API pricing is $10 per million input tokens and $50 per million output tokens, with separate cache rates. Fast mode can run at up to twice Standard speed for twice the applicable price.

Benchmark table comparing Muse Spark 1.3, Muse Spark 1.2, GPT-5.6 Sol, and Claude Opus 5 across agent, long-context, and coding tests. Muse Spark 1.3 leads both MRCR long-context benchmarks, DeepSWE v1.1, and SWEAtlas CodeBase QnA, while tying GPT-5.6 Sol at 88.8 on Terminal-Bench 2.1.

Meta released Muse Spark 1.3 for agentic and coding work, with a focus on sustaining longer tasks and managing several workflows inside one conversation. It can gather context with tools, recognize missing information in its plan, and preserve instructions as a task develops.

The model is also designed to collaborate more deliberately. It can ask clarifying questions when instructions are ambiguous, seek help when blocked, and confirm with the user before taking consequential actions.

For coding, Meta reports that Muse Spark 1.3 used about 20% fewer tool calls and 25% fewer tokens than Muse Spark 1.2 in comparisons conducted by its engineers. Meta also describes the model as taking fewer unnecessary turns and producing less verbose code.

Muse Spark 1.3 is available through Muse Code and the Meta Model API, including its max reasoning mode. Meta also reports stronger resistance to adversarial inputs and prompt injections.

Benchmark table showing Gemini 3.8 Flash leading several tests while keeping lower token pricing than most compared models.

Google released two versions of Gemini 3.8: Gemini 3.8 Flash for coding, reasoning, and agent workflows, and Gemini 3.8 Flash Cyber for vulnerability discovery and automated patching. Both share the same underlying model intelligence but use different safeguards and deployment rules.

Gemini 3.8 Flash scored 54.9% on HLE-Verified in Google’s testing. Google says the model can spend more reasoning steps and make additional tool calls on difficult tasks, while developers can lower the effort setting when reducing token use matters more.

For cybersecurity, Google reports that Flash Cyber exceeded 70% success on an internal vulnerability-discovery benchmark spanning 20 programming languages. On the external CWE-Bench patching benchmark, it scored 47.2%, compared with 47.8% for the leading frontier model in Google’s comparison.

Gemini 3.8 Flash costs $0.75 per million input tokens and $3.75 per million output tokens through December 31, 2026. Google says those rates rise to $1.50 and $7.50 on January 1, 2027. Developers can access it through the Gemini API and Google AI Studio, while Pro and Ultra subscribers can use it across several Google products.

Flash Cyber has more permissive cybersecurity safeguards and is restricted to trusted government authorities, critical-infrastructure operators, and software maintainers through Google’s Fairwind Program.

🧠RESEARCH

Researchers train web-agent world models to distinguish the true result of an action from outcomes produced by alternatives, rather than merely predict the next page state. Using branching WebArena trajectories, the method improved predicted-state matching, action ranking on WebPRMBench, and end-to-end success on WebArena-Lite, according to the authors’ experiments overall.

SafeEvolve improves agent safety by updating both the runtime harness and the underlying model policy from completed task experience. For Qwen3.5-4B on AgentDojo, the authors report a threefold reduction in attack success while benign-task utility increased from 59.79% to 61.86%, suggesting safety and usefulness can improve together in testing overall.

Repo-To-Skill turns machine-learning repositories into compact, verified instructions that research agents can reuse. Its AREX-Skill Library contains more than 5,000 skills from 1,000 repositories. With GPT-5.5 and the execution budget held fixed, the authors report gains across MLE-bench, PaperBench, FrontierCS, and PassNet compared with the same agent without skills overall.

📲SOCIAL MEDIA

🗞️MORE NEWS

Thinking Machines Lab is in talks to raise $5 billion to $6 billion at a pre-investment valuation of at least $40 billion, according to The Information. Nvidia is expected to contribute roughly half of the round, while Accel is in discussions to lead it; the financing remains open and its terms could change. Mira Murati-led startup generates at least a few hundred million dollars in annualized revenue. The proposed valuation remains below the $50 billion-plus valuation Thinking Machines discussed late last year.

Sen. Bernie Sanders and Rep. Greg Casar announced forthcoming legislation that would permanently prohibit the development and deployment of artificial superintelligence and temporarily pause advanced AI development until a new federal regulator establishes safety rules and a model-review process.

The proposal would create a cabinet-level AI agency with authority to monitor frontier systems and supervise the removal of dangerous capabilities. Its summary proposes corporate dissolution for entities that violate the restrictions and prison terms of up to 20 years for individuals, while directing the U.S. to pursue international restrictions as well. This is forthcoming legislation, not enacted law.

Nscale signed a multi-year agreement to provide Figure with an initial $3.5 billion in computing resources for its AI models and humanoid robots. Initial GPUs are targeted for deployment in Barstow, Texas, beginning in the second half of 2027. The companies intend to expand the commitment beyond $6 billion, with the potential deployment of up to 100,000 Nvidia GPUs. Nscale will also become a Figure shareholder and its preferred computing provider for future Helix models and humanoid robotics.

Moonshot AI has confidentially filed for a Hong Kong initial public offering and is aiming to raise $3 billion, according to Reuters sources. The Kimi developer was valued at roughly $50 billion in an ongoing funding round, although the IPO timetable and financial terms remain subject to regulatory approval and market conditions.

What'd you think of today's edition?

Login or Subscribe to participate in polls.

Reply

or to participate.