- NATURAL 20
- Posts
- AI Gets Cheaper, Smarter and More Useful Across Coding, Commerce and Agents
AI Gets Cheaper, Smarter and More Useful Across Coding, Commerce and Agents
PLUS: Xiaomi releases MiMo V2.6 and OpenAI forms an independent math advisory group.

AI agents built to get work done
Skydive agents take on work for you, using the same tools your team already uses.
Talk to your agent on the web, Slack, email, iMessage, or right from your terminal. Wherever you pick up the conversation, your agent keeps the context.
Put agents to work across customer support, sales, marketing, engineering, ops, and more.
Give your agent a job. They’ll take it from there.
Today:
Grok 4.7 targets longer coding and knowledge work at Grok 4.6 pricing
Claude Opus 5.5 brings Fable-level performance at a lower cost
Shopify brings Shop Pay checkout to Meta Muse
MiMo V2.6 scales reinforcement learning and opens its model weights
OpenAI is reportedly developing more persistent agent features for ChatGPT and Codex
Frontier Models Get Cheaper as AI Moves Deeper Into Everyday Work
Two frontier-model launches and a major commerce integration lead this edition. Grok 4.7 is pushing harder on multi-hour coding and professional work, Claude Opus 5.5 is bringing higher-end performance into a cheaper Opus tier, and Shopify is making Meta’s Muse a direct path to checkout.
The common theme is execution. The newest systems are being built to stay on tasks longer, use tools more reliably, and move from answering questions into workflows that end with code, research, or a completed purchase.
Grok 4.7 is xAI’s new frontier model for coding, agentic tasks, and professional knowledge work. It uses a larger base model than Grok 4.6 and received a longer reinforcement-learning run weighted toward harder tasks that can take hours to complete.
The gains show up most clearly in longer-running work. xAI reports 46.3% on CursorBench 4.0, up from 40.4% for Grok 4.6, along with 71.0% on DeepSWE v1.1 and 38.0% on Terminal-Bench 4.0; it also improved on professional-work benchmarks covering legal, clinical, and electrical-engineering tasks.
Grok 4.7 keeps the standard model at $2 per million input tokens and $6 per million output tokens for prompts below 200,000 tokens, with a 500,000-token context window. It is available through the xAI API, Grok Build, Cursor, model gateways, and a US regional endpoint, while a faster variant runs at roughly twice the output speed for twice the standard token price.
xAI also rebuilt the safeguard stack for the release. The company says Grok 4.7 improved jailbreak resistance and risky cyber and biology handling, while select cybersecurity partners are getting invite-only access to stronger red-team capabilities for defensive research.

Claude Opus 5.5 is the first model in Anthropic’s new Claude 5.5 family. Anthropic says it performs at the level of Claude Fable 5.1 on most work while costing about 40% less to run than Opus 5 because it uses fewer tokens per task and carries lower token prices.
The API price is $4 per million input tokens and $20 per million output tokens, 20% below Opus 5’s headline rates. Anthropic also says output is more than 30% faster, while its launch benchmarks put Opus 5.5 at 66.4% on Terminal-Bench 4.0, 57.8% on CursorBench 4.0, and 1,846 Elo on GDPval-AA v2.1.
Safety is a major part of the release. Anthropic says Opus 5.5 showed 85% fewer attempts to get around containment boundaries than Opus 5 or Mythos 5.1 in its internal testing, and outside evaluators including Frontier Design and METR tested the model before launch.
Claude Opus 5.5 is available across Claude products, the API, and major cloud platforms. Anthropic says Sonnet 5.5 and Haiku 5.5 will follow as it expands the new family.

Shopify is partnering with Meta to bring agentic checkout through Shop Pay into Muse across Shopify-powered stores. The goal is to let people move from a product conversation inside Muse to checkout without rebuilding the purchase flow somewhere else.
Shopify had already added Meta as an AI channel inside Agentic Storefronts. Products are shared with Meta by default through Shopify Catalog, while merchants can manage Catalog access, direct checkout, and Meta performance from their agentic settings.
The new checkout partnership turns that catalog connection into a transaction path. Muse can help users discover products from Shopify merchants, while Shop Pay handles the checkout layer across participating stores.
This puts Shopify’s commerce infrastructure behind another major AI shopping surface. Rather than requiring shoppers to begin on a merchant website, Shopify is positioning its catalog and checkout tools to work wherever AI agents are helping people decide what to buy.
🧠RESEARCH
RRSI tests whether AI agents can improve their own prompts, tools, memory, and control flow without overfitting to the tasks used for optimization. Across eight benchmarks, the method gains up to 14.1 points in-distribution and 4.7 points out-of-distribution while using 30% fewer policy tokens than unregularized evolution overall in testing.
GameHorizon introduces a large-scale benchmark for testing AI agents across short and long gameplay tasks. It combines 5,000 hours of recordings from 21 AAA games, 100 expert players, automated instruction annotation, and offline plus stepwise online evaluation. The authors test 47 models through more than one million model invocations overall.
WorldCrafter is a video world model designed to remember scenes across long camera movements and changing viewpoints. Its implicit 3D-aware memory compresses past observations into view-specific tokens before generation. Experiments on static and dynamic scenes report better long-horizon consistency and camera-control accuracy while preserving visual quality during minute-scale exploration overall.
📲SOCIAL MEDIA
🗞️MORE NEWS
Xiaomi released MiMo-V2.6-Pro and Flash as native multimodal models built for long-horizon agents, coding, computer use, design, research, and other professional workflows, with Pro reaching 46 points on the Artificial Analysis Intelligence Index in Xiaomi’s launch material. The company says a six-day reinforcement-learning run used roughly 750,000 trajectories, cost about $2.62 million for Pro and $850,000 for Flash, and raised DeepSWE v1.1 from 58.4 to 72.6 for Pro and 48.8 to 65.7 for Flash. The weights, training resources, and RL framework are being released openly, while the models keep V2.5 API pricing and support a 1-million-token context window.
OpenAI is reportedly developing features that would turn existing Codex and ChatGPT agent technology into more persistent assistants capable of handling routine tasks, managing software, and collaborating in the background. The company has also discussed a consumer personal assistant aimed at the same category as Meta’s Muse, although no launch timing has been announced. The work would extend technology OpenAI already has rather than requiring an entirely separate agent stack.
DeepSeek is reportedly making Chinese-made training hardware a major priority as access to advanced Nvidia chips remains constrained by US export controls. CEO Liang Wenfeng told investors that Huawei could begin delivering a new batch of training-capable chips as early as the fourth quarter, while DeepSeek is investing about 30 billion yuan in additional computing capacity and still relies partly on Nvidia hardware. The company is training a roughly 2-trillion-parameter model and has discussed eventually scaling toward 8 trillion parameters, making chip supply a central constraint on its roadmap.
OpenAI and Anthropic reportedly discussed a legally binding agreement that would let each company stress-test the other’s commercially available models for vulnerabilities and hidden risks. The proposed arrangement involved reciprocal API access, excluded unreleased models, and included commitments not to retain the other company’s testing data, but it remains unclear whether the agreement was finalized. The talks would turn one-off cross-lab evaluations into a more formal peer-testing process between two direct competitors.
OpenAI has formed an independent Advisory Group on Mathematics and Artificial Intelligence after saying an internal model trained since August 28 resolved the Navier–Stokes Millennium Prize problem and more than 100 other long-standing open problems, claims that still require external mathematical verification. The nine-member group is hosted at the Institute for Advanced Study, is unpaid by OpenAI, and can publish unsolicited advice or criticism while helping review significance, communication, and release plans for AI-generated results. The group will not control the pace of OpenAI’s internal mathematics research, leaving its role focused on review, context, and how results reach the wider community.
What'd you think of today's edition? |



Reply