- NATURAL 20
- Posts
- AI Gets Better at Voice, Images and Long-Running Agent Work
AI Gets Better at Voice, Images and Long-Running Agent Work
PLUS: Gemini crosses testing boundaries, and Anthropic opens a wet lab.

Blu Dot surpasses 2,000% ROAS with self-serve CTV ads
Home furniture brand Blu Dot blew up on CTV with help from Roku Ads Manager. Here’s how:
After a test campaign reached 211,000 households and achieved 1,010% ROAS, the brand went all in to promote its annual sales event. It removed age and income constraints to expand reach and shifted budget to custom audiences and retargeting, where intent was strongest.
The results speak for themselves. As Blu Dot increased their investment by 10x, ROAS jumped to 2,308% and more page-view conversions surpassed 50,000.
“For CTV campaigns, Roku has been a top performer,” said Claire Folkestad, Paid Media Strategist, Blu Dot. “Comping to our other platforms, we have seen really strong ROAS… and highly efficient CPMs, lower than any other CTV partner we've worked with.”
Using Roku Ads Manager, the campaign moved from a pilot to a permanent performance engine for the brand.
Today:
Grok Voice Transcribe 2.0 cuts transcription errors at the same price
Qwen Image 2.1 unifies generation, editing, and transparent images in a 7B model
StepFun Step 5 Preview brings 600B MoE scale to long-running agent work
Gemini crossed into three real companies during a cyber evaluation
Researchers used Claude during an authorized security test that exposed two critical flaws
Voice, Image, and Agent Models All Moved Forward This Weekend
Three different parts of the AI stack received meaningful upgrades over the weekend: speech recognition, image generation and editing, and long-running agent work.
Grok Voice Transcribe 2.0 targets messy real-world audio, Qwen-Image-2.1 brings broad editing and transparency into one compact open model, and Step 5 Preview is built to keep working across longer software and professional tasks.
Grok Voice Transcribe 2.0 is xAI’s new speech-to-text model for batch and real-time transcription. xAI says it is twice as accurate as Grok Voice Transcribe 1.0 across its real-world evaluations while keeping the same API pricing.
The model currently ranks first for accuracy among 32 streaming speech-to-text models on the Artificial Analysis leaderboard. On xAI’s short multilingual voice-command set covering 19 languages, word error rate fell from 20.6% with Transcribe 1.0 to 6.8% with Transcribe 2.0.
Pricing remains $0.10 per hour for batch transcription and $0.20 per hour for streaming. The API also includes word-level timestamps, speaker diarization at no additional cost, transcription of up to eight channels, biasing for as many as 100 key terms, automatic text formatting, filler-word removal, and turn detection for voice agents.
The model was trained for noisy, multilingual, real-world audio rather than only clean recordings. xAI says the underlying Grok Voice stack already handles tens of thousands of customer-support calls each day, millions of hours of video narration, and voice interactions in products including Tesla vehicles.

Qwen-Image-2.1 is a new open-source image model that combines text-to-image generation and image editing in one system. Its visual generation component has 7 billion parameters, making the release relatively compact while still targeting high-quality generation and detailed edits.
One of the biggest additions is native transparency. The model can generate regular or transparent images directly from a prompt, edit subjects while preserving transparent backgrounds, modify text inside transparent layers, and extract objects from ordinary photos as reusable RGBA assets.
Editing now supports up to 10 reference images at once. Users can mark local changes with circles, painted regions, or separate masks, while the model is designed to preserve facial identity, product shape, textures, and text across edits.
The model also supports panoramas, infographics, and storyboards, with improvements to typography, portrait lighting, and fine detail. Qwen uses a 32-layer Single-Stream DiT design plus mixed-granularity attention and KV-cache reuse to reduce the cost of working with multiple input images.

Step 5 Preview is StepFun’s new flagship model for agentic work, with a sparse Mixture-of-Experts architecture containing 600 billion total parameters but activating 27 billion per token. It supports a 1-million-token context window and vision input, with software engineering, professional knowledge work, and finance among its main targets.
StepFun reports a score of 67.7% on DeepSWE v1.1, 49.0% on StepCodeBench, and 66.4% on FrontierFinance. The model also scores 44 on the Artificial Analysis Intelligence Index.
The release puts extra emphasis on work that takes many steps. In one 24-hour experiment, Step 5 Preview repeatedly rewrote and benchmarked an inference kernel until it reached a reported 508 TFLOPS, while another experiment had the model improve a post-training data loop for a Qwen base model from 53.3% to 60% on AIME24.
Step 5 Preview is available now through StepFun’s products and API. StepFun says the model’s open weights will be released on October 15; most benchmark figures in the launch material are company-reported, while the Artificial Analysis index is externally measured.
🧠RESEARCH
DeepSeek-V4.1-Flash targets the memory and bandwidth costs of long-context agents with aggressive KV-cache compression. The 552B-parameter multimodal MoE activates 16B parameters during decoding and 8B during prefill, supports one-million-token contexts, cuts global KV cache to 890 bytes per token, and was pretrained on 45 trillion multimodal tokens before post-training began.
Researchers test whether MiniMax-H3 can reason about physical events when text, images, video, and audio each provide only part of the evidence. Across 517 cases, the model reaches 41.97% overall measured success. Video-based decision reasoning scores highest at 56.00%, while audio-based disambiguation is weakest at 27.40%, exposing multimodal integration limits.
This study finds that mismatched end-of-sequence tokens can make distilled language models generate unnecessarily long answers. Across Qwen3, Llama, and Gemma, treating equivalent stop tokens as one semantic action reduces length inflation. Training-stage analysis shows the mismatch is important but not the only cause, with late-stage inflation remaining after corrections.
📲SOCIAL MEDIA
🗞️MORE NEWS
During a May cybersecurity evaluation, a Gemini model accessed the systems of three real companies it incorrectly believed were part of the test. The model stopped once it recognized that the targets were real, and the affected companies were notified. Google confirmed the incidents, which are the first publicly known cases of one of its AI systems crossing into real external networks during testing.
A three-person security team used Claude during an authorized bug-bounty investigation and found two critical vulnerabilities affecting OpenAI systems. The researchers reached multiple employee ChatGPT accounts and internal software information, and OpenAI paid a $6,500 bounty after the findings were reported. OpenAI says the issues were fixed, while the case shows how capable AI tools can speed up legitimate security research and raise the importance of tightly controlled testing.
Naive AI, a Beijing startup founded in February by Tsinghua professor Jifeng Dai, has reportedly raised about $400 million across three rounds and reached a valuation above $1.4 billion. The company has fewer than 100 employees and is preparing an open-weight LLM that could arrive as soon as this month. Instead of pretraining from scratch, Naive is reportedly building on an existing Chinese open-weight model and concentrating on midtraining, post-training, reinforcement learning, and architectural changes.
Anthropic and Accenture are creating an embedded evaluation program that will place independent evaluators inside Anthropic with access comparable to employees. Faculty, Accenture’s specialist AI business, will red-team models, run alignment assessments, and test safeguards using knowledge of how enterprises and governments deploy AI in practice. Both companies expect to invest at least $1 billion each over the next five years to build this evaluation capacity, while the operating details are still being developed.
Anthropic has set up a wet lab in the San Francisco Bay Area to connect Claude-based research with physical biology experiments and automation. The company is hiring biology and biochemistry specialists and is expanding its life-sciences work after developing Claude Science and acquiring Coefficient Bio, with rare and neglected diseases among the areas of interest. Anthropic says the lab is not solely for drug discovery and does not plan to enter commercial clinical trials, keeping its role focused on research and tools rather than becoming a drug company.
Anthropic has reportedly moved its planned IPO from October to November, allowing the company to include third-quarter financial results in its investor pitch. The offering could seek as much as $100 billion at a valuation around $2 trillion, although the size, valuation, and timing remain subject to change before an offering is completed. The shift was reportedly decided before the latest public debate over frontier-AI safety, separating the scheduling change from those more recent concerns.
What'd you think of today's edition? |



Reply