• NATURAL 20
  • Posts
  • AI Costs Keep Falling as GPT-6 and New Voice Models Roll Out

AI Costs Keep Falling as GPT-6 and New Voice Models Roll Out

PLUS: Claude finds a new enzyme system, and OpenAI launches MentalHealthBench.

In partnership with

Is Your Training Data Actually Model-Ready?

If you're fine-tuning a speech model, you've probably hit this wall: DNSMOS gives you a score, but it doesn't tell you whether the data behind that score is actually right for your model. 

Treat it as a pass/fail gate and you'll end up training on audio that looks clean on paper but drags down real-world performance—while good source data gets tossed for no reason.

Voices' CTO DJ Jalali (with the team's senior audio and voice data engineers) just published a free white paper that breaks down the four-step calibration framework they use internally to set model-specific quality thresholds instead of trusting the raw DNSMOS number. It also covers where DNSMOS breaks down and how Voices validates audio for custom datasets at scale.

Today:

  • GPT-6 Sol and Luna bring GPT-6 into lower-cost everyday work

  • Gemini 3.8 TTS adds custom voices, voice replication, and cheaper audio generation

  • Qwen Audio 3.1 launches five voice models and cuts API prices by up to 95%

  • ChatGPT Voice can now run on GPT-6 models and use connected apps

  • Scribe v2 Medical cuts transcription errors on clinical audio

GPT-6 Gets Cheaper While Voice AI Becomes More Customizable

This edition is led by a sharp drop in the cost of frontier reasoning and two major upgrades to voice generation. OpenAI is bringing GPT-6 capabilities into cheaper Sol and Luna models, while Google and Qwen are pushing voice systems toward custom characters, lower latency, and much lower API costs.

The result is a broader shift in how these systems are being packaged. High-end reasoning is getting cheaper, while audio models are moving beyond basic speech synthesis into voice design, replication, real-time interaction, and specialized production workflows.

AutomationBench chart comparing GPT-6 Astra, Sol, and Luna with GPT-5.6 and Claude models, showing benchmark scores against cost per task.

GPT-6 Sol and GPT-6 Luna extend the GPT-6 family below Astra, giving developers and ChatGPT users two cheaper options built around the same generation of reasoning models. Sol is aimed at complex coding and agentic workflows, while Luna is designed for focused, high-volume work where cost and speed matter more.

The biggest change is price. Sol costs $2 per million input tokens, $0.20 for cached input, and $10 per million output tokens, while Luna costs $0.10 per million input tokens, $0.01 cached, and $0.50 per million output tokens—roughly 50% below the promotional GPT-5.6 rates OpenAI had been offering.

Both models support text and image input, a 1.05-million-token context window, up to 128,000 output tokens, and reasoning settings from none through max. OpenAI positions Sol between Astra and Luna for demanding work, while Luna brings GPT-6-level tooling to workloads that previously would have been too expensive to run at scale.

Sol and Luna are available through the API, ChatGPT Work, and Codex. Plus, Pro, Business, Enterprise, and Edu users can access them in Work and Codex, while Free and Go users can try Luna in the desktop app.

Text-to-speech benchmark comparing Gemini 3.8 Flash TTS and Flash-Lite TTS with Gemini 3.1, ElevenLabs, Cartesia, OpenAI, and Inworld across overall quality, human-like variation, multi-speaker performance, and style control.

Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS are Google’s new text-to-speech models for expressive audio generation and high-volume voice applications. Flash is the higher-quality creative tier, while Flash-Lite is optimized for lower-cost, lower-latency production and real-time voice-agent pipelines.

The models can create voices from natural-language descriptions, replicate existing voices with consent verification, and work with an extended library of more than 150 prebuilt and custom voices. Users can also direct dialogue line by line, controlling pacing, emotion, pronunciation, conversational sounds, and character delivery.

Pricing is $0.50 per million text input tokens for both models through the end of 2026. Audio output costs $9 per million tokens for Flash TTS and $6 for Flash-Lite TTS, with Google also offering batch pricing at half those output rates; the standard prices are scheduled to double on January 1, 2027.

The models are available across Google AI Studio, the Gemini API, Gemini Enterprise, Gemini Notebook, and Google Vids. Google is positioning Flash for audiobooks, podcasts, character work, and other high-fidelity production, while Flash-Lite targets large-scale and real-time voice applications.

Promotional graphic for Qwen Audio 3.1 highlighting speech recognition, text-to-speech, and real-time conversational audio capabilities.

Qwen Audio 3.1 expands Alibaba’s audio stack with five models covering speech recognition, text-to-speech, real-time conversation, audio understanding, and expressive voice generation. The lineup includes upgraded ASR, TTS, and Realtime models alongside the new ASR-Next and TTS-Next systems.

Qwen says the ASR model reaches an average character error rate of 4.55% across its public dialect tests, while the Realtime model can listen while speaking, handle interruptions, and change its response style based on signals such as a user sounding upset. TTS-Next is aimed at more expressive audio creation rather than simple read-aloud speech.

The release also comes with major API price cuts. Qwen says TTS pricing is down about 70%, Realtime pricing about 85%, and ASR pricing by as much as 95%, with most of the new models already available through its AI platform.

Together, the five-model lineup gives Qwen separate tools for recognition, synthesis, live interaction, deeper audio understanding, and creative speech generation instead of forcing all audio workloads through one general-purpose model.

🧠RESEARCH

Taste-Bench measures whether AI agents choose the better path at important decision forks during long tasks. Across 502 forks mined from engineering and research trajectories, the best frontier model reaches 59.7%. Training a student on hindsight-informed judgments improves unseen-task decisions and raises SWE-bench Pro success from 14.6% to 33.7% overall.

WorldCrafter gives video world models an implicit 3D-aware memory so they can revisit scenes more consistently across long camera movements. Instead of storing explicit geometry, it compresses past observations into viewpoint-specific tokens. Experiments report better long-horizon consistency, stronger camera control, and preserved visual quality during minute-scale exploration from one prompt.

Designer-RSI lets a frozen AI design agent improve by evolving an external memory of reusable procedures instead of updating model weights. Across five rounds, its skill bank grows from 76 to 139 procedures. GenEval2 execution success rises from 72.7% to 99.3%, using real user briefs directly without human reward labels.

📲SOCIAL MEDIA

🗞️MORE NEWS

ChatGPT Voice can now use GPT-6 Astra, Sol, and Luna, extending the new model family into spoken conversations. Voice can also work with connected apps such as email, calendars, and Slack, including inside ChatGPT Work. The update lets users combine live conversation with the same external context and tools already available in text-based workflows.

Scribe v2 Medical is now generally available as a speech-recognition model tuned for clinical dictation, medication names, dosages, anatomy, and other medical vocabulary. ElevenLabs says the medical fine-tune reduces word error rate by 18% compared with its base Scribe v2 model and achieved its lowest reported error rate on reproducible medical ASR benchmarks. It is available through ElevenAPI at the same rate as Scribe v2 and keeps features including keyterm prompting, entity detection, speaker diarization, and no-verbatim mode.

Claude identified a previously uncharacterized enzyme system associated with long arrays of evenly spaced DNA repeats while analyzing biological sequence data with high-level direction from researchers. The system, called ART, appears mainly in bacteriophages and includes a reverse transcriptase, a neighboring partner gene, and repeat sequences that early experiments show are expressed as distinct short RNAs. Anthropic says more experiments are underway to determine how the system works, so the finding is an early biological discovery rather than a fully characterized mechanism.

MentalHealthBench is a new open benchmark built with more than 80 licensed psychologists and psychiatrists from 22 countries to test AI responses in realistic mental-health conversations. It contains 1,215 synthetic conversations and 5,262 expert-written rubric criteria spanning everyday non-acute situations, high-acuity concerns, emergencies, adults, teens, caregivers, clinicians, and 19 languages. The benchmark measures behaviors such as safety, context seeking, practical guidance, and preservation of user agency, with OpenAI releasing the data so other researchers can run and inspect the evaluation.

President Donald Trump said during his September 22 United Nations General Assembly speech that he was directing US government documents to replace the phrase artificial intelligence with super intelligence. He argued that the word artificial makes the technology sound fake and paired the proposed terminology change with a broader call to encourage AI development rather than slow it, while also saying the technology requires care. The term superintelligence already has an established technical meaning for systems that exceed human intelligence, and it remains unclear whether the proposed wording change will be implemented across federal agencies.

What'd you think of today's edition?

Login or Subscribe to participate in polls.

Reply

or to participate.