• NATURAL 20
  • Posts
  • OpenAI Astra Tackles 10 Open Problems as DeepSeek and ByteDance Ship Major Updates

OpenAI Astra Tackles 10 Open Problems as DeepSeek and ByteDance Ship Major Updates

PLUS: Google pauses AI image generation in Earth, while Chinese military researchers use U.S. model outputs to train local defense systems.

In partnership with

Your prompts are leaving out 80% of what you're thinking.

When you type a prompt, you summarize. When you speak one, you explain. Wispr Flow captures your full reasoning — constraints, edge cases, examples, tone — and turns it into clean, structured text you paste into ChatGPT, Claude, or any AI tool. The difference shows up immediately. More context in, fewer follow-ups out.

89% of messages sent with zero edits. Used by teams at OpenAI, Vercel, and Clay. Try Wispr Flow free — works on Mac, Windows, and iPhone.

Today:

  • Astra produces ten advances on long-standing open problems

  • Seedance 2.5 doubles single-pass video length to 30 seconds

  • V4-Flash API beta adds stronger agent performance and Codex support

  • AI image generation in Earth is paused after policy violations

  • Chinese military researchers use U.S. AI outputs to train local systems

AI Is Moving From Benchmark Scores to Real Work

The most important AI releases this weekend were not ordinary chatbot upgrades. OpenAI showed an unreleased model contributing to mathematical research, ByteDance pushed AI video closer to a complete production workflow, and DeepSeek made its low-cost model more capable of operating coding tools.

Together, the stories point to a larger shift: the next competitive advantage will come from models that can discover, create, and act not merely answer questions.

OpenAI

OpenAI says an internal version of Astra, its next major model family, produced new results for ten problems across geometry, coding theory, group theory, quantum complexity, cryptography, and combinatorics. The problems had seen no progress on their main result for at least a decade, and OpenAI estimates that the tokens used to find all ten solutions would cost roughly $2,000 at GPT-5.6 Sol API rates.

Humans used the model to prepare the arguments as manuscripts, after which Astra formalized each proof in Lean so the logic could be checked by software. The most striking results include a construction for non-sofic groups, new sphere-packing bounds, and advances tied to post-quantum cryptography.

This is more meaningful than another benchmark win because the work targets unanswered research questions. Astra remains unreleased, however, and the mathematical community still needs to independently examine whether every formal statement correctly represents the original problem.

ByteDance

Seedance 2.5 can generate synchronized audio-video clips up to 30 seconds long in one pass, twice Seedance 2.0’s limit, then extend them through additional rounds. It can organize several connected shots inside one generation rather than simply stretching a single scene.

The model accepts as many as 30 images, 10 video clips, and 10 audio clips at once. It also adds timestamp-level editing, more stable character and scene continuity, and finer control over green screens, camera perspective, motion, and reference-based changes.

Dreamina says its international rollout has begun for eligible subscriber accounts, while ByteDance says API access will come later through BytePlus ModelArk. That staged availability matters: this is a product launch, but access may not appear for every region or account at the same time.

DeepSeek

DeepSeek has moved the official V4-Flash API into public beta after a new post-training stage focused on coding and agent tasks. Company-reported results include 82.7 on Terminal Bench 2.1, 54.4 on DeepSWE, and 70.3 on Toolathlon Verified—higher than DeepSeek’s earlier V4-Pro preview across its listed agent tests.

The model now supports the Responses API natively and can be configured for use in Codex. Its architecture and size are unchanged from V4-Flash-Preview; the gains come from a new post-training stage rather than a larger model.

V4-Flash supports a one-million-token context window and currently costs $0.14 per million uncached input tokens and $0.28 per million output tokens. The update applies only to the API: V4-Pro and DeepSeek’s app and web models remain unchanged, and the benchmark figures are not independent test results.

🧠RESEARCH

ReToken trains one special embedding to retrieve the visual tokens most relevant to a question, avoiding the cost of processing every frame or image at once. The authors report gains of 13.4 points for Qwen3VL-8B on Visual Haystacks and 8.0 points on the long-video LVBench test, with training and inference fitting on a single H100 GPU.

PAC-MAN combines reinforcement learning with safety constraints so a Unitree G1 humanoid can protect its entire body using only a head-mounted depth camera and its internal sensors. The real robot avoided contact on 95% of test throws, showing how safety rules can be trained around imperfect real-world perception instead of relying on exact object tracking.

Sample More, Reflect Less compared seven reasoning methods across 1.5B, 3B, and 7B open models while counting every generated token. None of the 36 method comparisons reliably beat simply generating several answers and choosing the majority result at the same cost; ten were significantly worse, suggesting some apparent “self-reflection” gains may come mostly from spending more tokens.

📲SOCIAL MEDIA

🗞️MORE NEWS

Reuters reports that Google paused a Nano Banana 2 feature that let users generate photorealistic scenes grounded in Google Earth imagery, one day after launch. Users had shared outputs that appeared to violate company policies, raising concerns about fake events placed over recognizable real-world locations; Google says the images were watermarked and did not appear in the main Earth experience.

A Reuters review of more than 80 Chinese papers and patents found military-linked researchers using outputs from OpenAI and Anthropic models to train smaller systems for surveillance, cyber operations, code analysis, and tactical uses. This process, called distillation, can transfer selected capabilities into models that run locally, but it cannot reproduce the full ability of a frontier model or replace the computing needed to train one.

MiniMax’s new H3 model can generate clips up to 15 seconds in 2K with native stereo audio and can use text, images, video, and audio as inputs. The company says it will release the model weights, support several Chinese-made chips, and price 2K generation at less than one-third the cost of mainstream rivals; those price and performance claims still need outside testing.

A Munich court ruled that AI music company Suno lacked the rights to process songs represented by licensing group GEMA and ordered it to disclose revenue connected to the infringement. Damages have not yet been set, and Suno says it may appeal, but the ruling adds pressure on generative-AI companies to license copyrighted training material.

What'd you think of today's edition?

Login or Subscribe to participate in polls.

Reply

or to participate.