- NATURAL 20
- Posts
- OpenAI Wants Voice AI to Think, Translate, and Act in Real Time
OpenAI Wants Voice AI to Think, Translate, and Act in Real Time
PLUS: EVE Online Enters a New Era as Its Studio Goes Independent Again, Claude Agents Learn to Dream, Plan, and Work as a Team and more.

Become An AI Expert In Just 5 Minutes
If you’re a decision maker at your company, you need to be on the bleeding edge of, well, everything. But before you go signing up for seminars, conferences, lunch ‘n learns, and all that jazz, just know there’s a far better (and simpler) way: Subscribing to The Deep View.
This daily newsletter condenses everything you need to know about the latest and greatest AI developments into a 5-minute read. Squeeze it into your morning coffee break and before you know it, you’ll be an expert too.
Subscribe right here. It’s totally free, wildly informative, and trusted by 600,000+ readers at Google, Meta, Microsoft, and beyond.
Today:
OpenAI Wants Voice AI to Think, Translate, and Act in Real Time
Anthropic’s SpaceX Deal Gives Claude a Bigger Compute Engine
Google’s Gemini Flash-Lite Goes Live for Fast, Low-Cost AI Apps
EVE Online Enters a New Era as Its Studio Goes Independent Again
Claude Agents Learn to Dream, Plan, and Work as a Team

OpenAI launched three new audio models for developers: GPT-Realtime-2, GPT-Realtime-Translate, and GPT-Realtime-Whisper. The big idea is simple: voice AI should not just answer questions — it should listen, reason, translate, transcribe, call tools, and take action while a conversation is happening.
GPT-Realtime-2 is the main upgrade. It brings GPT-5-class reasoning to live voice apps, supports parallel tool calls, handles interruptions better, has a larger 128K context window, and lets developers adjust how much reasoning the model uses depending on the task. OpenAI says it performs better than the previous realtime model on audio reasoning and instruction-following tests.
The translation model supports 70+ input languages and 13 output languages, making it useful for customer support, travel, events, education, and global sales. GPT-Realtime-Whisper adds live speech-to-text, so captions, meeting notes, support workflows, and voice agents can work as someone is speaking.
GPT-Realtime-2 costs $32 per 1M audio input tokens and $64 per 1M audio output tokens. GPT-Realtime-Translate costs $0.034 per minute, while GPT-Realtime-Whisper costs $0.017 per minute.
Anthropic announced a major compute deal with SpaceX that gives it access to all compute capacity at SpaceX’s Colossus 1 data center. That means more than 300 megawatts of new capacity and over 220,000 NVIDIA GPUs coming online within the month.
Claude users get higher limits. Anthropic is doubling Claude Code’s five-hour rate limits for Pro, Max, Team, and seat-based Enterprise plans. It is also removing peak-hour limit reductions for Pro and Max users, while raising API rate limits for Claude Opus models.
This is part of Anthropic’s much bigger infrastructure race. The company also pointed to deals with Amazon, Google and Broadcom, Microsoft and NVIDIA, and Fluidstack, including multi-gigawatt agreements and a $30 billion Azure capacity partnership.
Anthropic says it is also interested in working with SpaceX on orbital AI compute, meaning AI data centers in space. That shows how far the compute race is stretching as AI companies search for more power, more chips, and more places to run models.

Google made Gemini 3.1 Flash-Lite generally available on the Gemini Enterprise Agent Platform. Google calls it its fastest and most cost-efficient Gemini 3 series model, built for low-latency tasks, high-volume apps, and cheaper large-scale AI workflows.
The model is aimed at companies that need AI to respond quickly without spending too much. Google says developers are using it for tool calling, agent coordination, automated pipelines, code assistance, customer support, gaming, creative tools, and financial workflows.
Several customer examples show why this matters. Gladly uses Flash-Lite for customer service across SMS, WhatsApp, and Instagram, reporting about 60% lower costs than comparable thinking-tier models, with roughly 1.8-second p95 latency for full replies and about 99.6% success rate under heavy load.
Google also highlighted use cases from JetBrains, Astrocade, Krea, OffDeal, Ramp, and AlphaSense. The common theme is speed: Flash-Lite is being positioned as the model for AI features that need to run constantly, respond quickly, and stay affordable at scale.
🧠RESEARCH
Google tested SymptomAI, a Fitbit-based chat agent that asks symptom questions and suggests possible causes, with 13,917 people. In 517 clinically reviewed cases, its possible-cause lists were more accurate than doctors reviewing the same chat. The study also linked symptoms with wearable signals, but warns the labels relied on self-reports.
Stream-R1 tries to make streaming video faster without losing quality. It improves distillation, meaning training a smaller model to copy a bigger one, by learning more from reliable examples and focusing on weak frames or image areas. Tests showed better quality, motion, and text matching at no extra use-time cost.
OpenSearch-VL is an open guide for building AI agents that answer image questions by searching, checking evidence, and using tools. It adds training data, visual tools like cropping and OCR, which reads text in images, and failure handling. Across seven tests, it gained 10+ points and matched closed commercial models.
📲SOCIAL MEDIA
🗞️MORE NEWS
EVE’s Maker Goes Independent Again The company behind EVE Online is becoming independent again under a new name, Fenris Creations, after years under Pearl Abyss. The EVE team says there will be no layoffs, no restructuring, and no change to its game plans, while also starting a research partnership with Google DeepMind to study AI in complex game worlds.
Claude Agents Get “Dreaming” Anthropic added dreaming to Claude Managed Agents, which means agents can review past work, spot patterns, and improve their memory over time. It also added outcomes, where developers define what “good” looks like, plus multi-agent orchestration, where one lead agent can split big tasks across smaller specialist agents.
Perplexity Opens Up Its AI Engine Room Perplexity explained how it uses CuTeDSL, a special coding language for NVIDIA AI chips, inside its in-house serving system called ROSE. The goal is faster AI responses: Perplexity says some custom kernels, or tiny chip-level programs, now run 2–3x faster, while other changes improved chip-to-chip data movement.
Perplexity Adds Finance Search for AI Agents Perplexity launched Finance Search in its Agent API, letting developers pull live market prices, company financials, earnings data, analyst estimates, and ETF details in one tool. This is meant for finance agents that need current numbers, not stale answers, such as stock research tools, earnings summaries, and company comparison apps.
ElevenLabs Agents Can Now Handle Files, Images, and Locations ElevenLabs expanded ElevenAgents beyond voice and chat so they can process images, PDFs, audio notes, contacts, and location pins. That means a support agent could read a customer’s uploaded document, inspect a photo, or continue the same conversation across voice, WhatsApp, and a web widget without making the user start over.
Google Brings AI Music Tools to Real Artists Google is partnering with Believe and TuneCore to bring Flow Music and Lyria 3 Pro to artists, producers, and songwriters. Flow Music acts like a creative assistant for lyrics, melodies, genres, and new instruments, and Google says it does not claim ownership of original content generated with the tool.
DeepSeek’s Value Could Jump to $45 Billion DeepSeek is in talks for its first venture funding round, with its possible valuation jumping from around $20 billion to $45 billion in just weeks. The round could help DeepSeek retain top researchers, while China’s chip fund and cloud giants like Tencent and Alibaba are reportedly interested.
What'd you think of today's edition? |


Reply