• NATURAL 20
  • Posts
  • Google Resets AI Leadership as Jeff Dean Leaves to Build Discovery Loop

Google Resets AI Leadership as Jeff Dean Leaves to Build Discovery Loop

PLUS: Anthropic starts designing Claude chips, and AMD buys Taalas to speed up AI inference.

In partnership with

Porkbun is the domain name registrar you need.

Still using GoDaddy or Namecheap? There’s a better way with Porkbun!

Porkbun is the domain registrar trusted by creators, developers, entrepreneurs, and folks who want low prices without the nonsense.

Why people are choosing Porkbun:
• Most domains sold at cost
• Low, transparent registration and renewal pricing
• Free features like WHOIS privacy and SSL certificates
• Powerful web and email hosting options
• Real human support 24/7, 365 days a year
• Named the #1 domain registrar by Forbes Advisor and USA Today

For launching a business, building a personal brand, starting a side project, or creating your first website, Porkbun makes it easy.

Get $1 off your next domain registration with Porkbun now.

Today:

  • DeepMind Leadership Resets as Jeff Dean Leaves for Discovery Loop

  • Muse Code Launches a Persistent Terminal Agent Powered by Muse Spark 1.2

  • GPT-5.6 Sol Gets More Focused Answers as Luna Expands to Free Users

  • Claude Chip Team Starts Building Custom Silicon

  • Taalas Acquisition Expands AI Inference Hardware Push

Big Picture

This week’s strongest signal is that AI competition is moving beyond model scores. Google is reorganizing the people who build frontier systems, Meta is packaging a model and agent runtime together, OpenAI is tuning how models behave for everyday users, and Anthropic and AMD are moving deeper into custom hardware.

The stack is becoming the product: model, agent harness, interface, chips, memory, and infrastructure increasingly have to work together. That makes execution speed and cost control as important as raw benchmark performance.

GOOGLE
Abstract Google AI leadership graphic showing a leadership reset and Discovery Loop launch.

Google is reorganizing its AI leadership while several of its best-known researchers leave to start a new company. Demis Hassabis is stepping back from day-to-day leadership of Google DeepMind to become Alphabet’s chief scientist and DeepMind chairman, while CTO Koray Kavukcuoglu takes a larger operating role.

At the same time, Jeff Dean, Sanjay Ghemawat, Oriol Vinyals, and Quoc Le are leaving Google to create Discovery Loop, a public-benefit research company focused on using AI to accelerate scientific discovery. The change separates more of Google’s frontier research leadership from the daily job of shipping Gemini products.

This is not a retreat from AI research. It is a split in responsibilities: Google is putting more operational control around Gemini while Hassabis focuses on longer-term science and AGI work. Discovery Loop also gives four veteran researchers a new vehicle to pursue automated science outside Google.

META
Bar charts comparing Meta Muse Code and Muse Spark 1.2 with leading coding models across Terminal-Bench 2.1, DeepSWE 1.1, and Meta’s internal coding benchmark.

Meta introduced Muse Code, a beta terminal coding agent powered by Muse Spark 1.2. It can plan software changes, write code, validate its work, and coordinate persistent subagents across large repositories.

Its runtime keeps a local event log of model calls, tool runs, approvals, and edits, making sessions replayable and restart-safe. Built-in skills include /plan for approval-gated planning, /grill for stress-testing a plan, and /goal for working toward a defined result.

Muse Spark 1.2 is trained more heavily on coding, debugging, repository understanding, and long-horizon workflows. Meta says it also co-trained the model with the Muse Code harness and tested kernel optimization runs exceeding 1,000 tool calls and lasting up to 24 hours.

Availability / pricing: Muse Code is in beta for macOS and Linux. Meta’s announcement provides installation instructions but does not list paid pricing in the post.

OPENAI
Side-by-side comparison showing GPT-5.5 Instant and the updated GPT-5.6 Sol answering a weather question, with GPT-5.6 giving a more direct recommendation and clearer rain and wind details.

OpenAI updated GPT-5.6 Sol in ChatGPT to give more direct answers, use tighter formatting, adapt detail to the question, and rely more carefully on sources for factual claims. Plus and Pro users also get a new slider for choosing how much reasoning ChatGPT should use.

OpenAI says an internal evaluation of financial, medical, and legal prompts found responses containing at least one factual error were about 68% less common with GPT-5.6 Sol and 62% less common with GPT-5.6 Luna than with GPT-5.5 Instant. Those are company-run evaluation results, not independent benchmarks.

Free access is expanding too. GPT-5.6 Luna will become the default for Free and Go users this week, with unlimited text chats and a new Think button arriving next week, subject to abuse guardrails. File uploads, images, and other tools will still have limits.

Availability / safety: The updated Sol and reasoning slider are available to Plus and Pro users now. OpenAI says the Chat-optimized Sol update does not change the version used in Work or Codex, and it added age-appropriate safeguards for users it believes are under 18.

🧠RESEARCH

Argus gives fixed-weight AI agents a persistent runtime with separate manager, planner, engineer, and reviewer roles. Across seven GPT-5.5 benchmark settings, it reached about 78% on SWE-Bench Pro versus 59% for a direct copilot. Later runs used fewer input tokens and less active time after verification-guided learning from earlier attempts.

ContextWeave tests whether remembered experience improves realistic, multi-month office workflows. Its 1,005 tasks reconstruct work from 14 participants with privacy protections. The strongest memory setup raised workspace quality from 68.08 to 78.20 and preference alignment from 41.50 to 70.60, while also showing that misleading memories can hurt execution reliability.

WorldCycle trains video world models using reversible action loops that should return to their starting state, creating supervision without labeled future frames. The method reduced state-return drift by up to 44% and nearly quadrupled composite-action accuracy over the base model. The authors also release CycleBench to test long-horizon physical consistency.

📲SOCIAL MEDIA

🗞️MORE NEWS

Anthropic is creating an in-house team to design custom chips for Claude and is hiring engineers across hardware and software. The company says custom silicon will complement, not replace, its multi-chip strategy using AWS, Google, Nvidia, and AMD hardware; it has not given a development or manufacturing timeline.

AMD is buying Toronto chip startup Taalas for an undisclosed amount to strengthen its AI inference technology. Taalas had raised about $219 million, and AMD plans to fold its specialized silicon and engineering team into an accelerator roadmap built around Instinct GPUs and broader system-level AI products.

DeepSeek has resumed a fundraising effort seeking nearly $8 billion at a reported valuation of about 500 billion yuan, or roughly $74 billion, according to Bloomberg reporting cited by Reuters. The round had been paused in July, so its restart points to renewed momentum around DeepSeek’s capital and possible IPO plans.

DeepSeek invested 140.8 million yuan, about $20.8 million, in Unitree’s Shanghai IPO for a 2.31% stake through a strategic placement. The companies plan to work together on embodied intelligence, combining DeepSeek’s AI expertise with Unitree’s robotics, motion-control, and physical-world data.

What'd you think of today's edition?

Login or Subscribe to participate in polls.

Reply

or to participate.