- NATURAL 20
- Posts
- AI Gets Faster, Cheaper and More Autonomous
AI Gets Faster, Cheaper and More Autonomous
PLUS: Grok 4.6 keeps its price, while ChatGPT adds cross-app Computer History.

[Webinar] Can you prove AI is working?
AI is in your engineering workflow. While the token spend shows it, the throughput doesn't. The human is very much still in the loop, and that's a context problem.
Join live on Aug 19 (FREE) to see:
The 4 metrics to measure the gap where gains leak out before production.
The 8 stages of context maturity, the specific walls capping your metrics, and a free tool to pinpoint where your team is
Why more MCPs and bigger context windows aren’t enough and what it takes to get real value from your agents.
Today:
Gemini 3.7 Flash targets coding and agents at half 3.6 Flash’s original price
V4 Pro launches with agent upgrades and flexible reasoning
Ultrafast runs GPT-5.6 Sol at up to 14× speed
Grok 4.6 improves on Grok 4.5 at the same price
ChatGPT Computer History remembers activity across apps and websites
AI Gets Faster, Cheaper and More Autonomous
The newest model releases are competing on a practical metric: how much high-quality work an AI system can complete before latency, token cost, or orchestration overhead gets in the way.
Google is pairing better coding and agent performance with temporary launch pricing, DeepSeek is tuning reasoning effort for different workloads, and OpenAI is pushing its strongest model toward real-time use.
The same week also shows why raw speed is only half the story. More persistent memory and multi-agent coordination can improve usefulness, but they also expand the privacy, reliability, and security questions teams must manage.
Google introduced Gemini 3.7 Flash as its most capable workhorse model yet for coding and agents, only three weeks after 3.6 Flash. The company says the update improves software engineering, knowledge work, web development, multi-step planning, and tool use.
On company-reported evaluations, 3.7 Flash scored 43.6% on FrontierCode 1.1 Main versus 34.4% for 3.6 Flash, and 65.3% versus 49.0% on DeepSWE v1.1. Google also reports gains in complex-document understanding and business-workflow automation.
The launch price is $0.75 per million input tokens and $3.75 per million output tokens through December 31, 2026—half the original 3.6 Flash price. Starting January 1, 2027, the rates are scheduled to rise to $1.50 and $7.50.
Developers can use the model through the Gemini API, Google AI Studio, Android Studio, and Antigravity. It is also rolling into Gemini Spark for Google AI Pro and Ultra subscribers in supported countries, while enterprises can access it through Gemini Enterprise.
DEEPSEEK

DeepSeek launched V4-Pro-0813, the production release of its flagship V4 Pro model, with what the company describes as major upgrades for agent workflows. The model is available through DeepSeek’s API, app, and web interface.
V4 Pro and V4 Flash now support adjustable reasoning effort: low for simple requests, high for everyday agent work, and max for complex tasks. DeepSeek also added native support for the OpenAI Responses API, reducing migration work for some agent developers.
The release turns reasoning depth into an operational control rather than a one-size-fits-all setting. Teams can reserve the most expensive thinking mode for hard tasks while using lighter settings where latency and cost matter more.
DeepSeek is also introducing peak and off-peak API pricing beginning August 17. The company’s performance improvements are self-reported, and the new model still needs independent testing across real production workloads.
OPENAI

OpenAI introduced Ultrafast, a new API service tier that runs GPT-5.6 Sol up to 14 times faster than Standard processing. The company says Cerebras hardware enables generation of as many as 750 output tokens per second.
The speed is aimed at work where waiting changes the result: incident response, fraud detection, live financial research, customer support, voice systems, commerce, and interactive experimentation.
OpenAI says its engineers use Ultrafast to review logs, traces, and conversations while outages are still unfolding, and to compress research loops that previously ran overnight into interactive work sessions.
Access remains a limited preview for selected customers. OpenAI has not published pricing, a general-release date, or independent latency tests, so the 14× and 750-token figures are company-reported maximums rather than guarantees for every workload.
🧠RESEARCH
VAKRA evaluates agents across more than 8,000 executable APIs in 62 domains. The best model reached 70.4% on simple endpoint tasks but only 50–51% on compositional APIs, with performance falling over 50% as reasoning deepened. Failures centered on disambiguation and cross-source grounding, not basic tool invocation mechanics alone in practice.
Researchers ran 56,476 inferences across four models, three reasoning benchmarks, and seven output limits from 64 to 4,096 tokens. Model rankings reversed at budgets on every benchmark, while 3–19% of questions became less accurate with more tokens. The result argues evaluations should report performance at several compute budgets, not one.
Researchers traced 307 agent failures to loaded skills across two benchmarks: 125 functional errors and 182 efficiency regressions. Seemingly relevant guidance often caused missing or incorrect implementation, while excessive verification created 67 cost regressions. The findings show reusable skills need comparisons, monitoring, and removal paths instead of trust in deployment.
🗞️MORE NEWS
SpaceXAI introduced Grok 4.6 as a significant improvement over Grok 4.5 while keeping the same price. The announcement did not provide a full benchmark table, so broad claims of frontier performance should remain provisional until independent evaluations cover accuracy, reliability, and cost.
OpenAI says Computer History in the ChatGPT desktop app can remember a user’s activity across apps and websites so later interactions require less explanation. The announcement did not spell out rollout scope, retention, or enterprise controls, making settings and privacy documentation important before teams enable it broadly.
Anthropic’s controlled experiments found that agent swarms can specialize and discover software vulnerabilities, but can also amplify conformity, collude in pricing games, or sabotage peers in shared environments. The results are laboratory studies rather than evidence that every deployed multi-agent system will behave this way, but they show single-agent testing is not enough.
Dream Research Labs says it recovered a 1,395-file workspace from a near-autonomous campaign that ran 12 attack waves over roughly four days, cracked 85 credentials, and exfiltrated more than 2,500 personnel records. The report links the framework to Hermes and OpenClaw agents; its Chinese-language operator assessment is an inference from artifacts, not a definitive state attribution.
What'd you think of today's edition? |



Reply