• NATURAL 20
  • Posts
  • New AI Models Target Better Reasoning, Search and Long-Running Work

New AI Models Target Better Reasoning, Search and Long-Running Work

PLUS: Claude moves into Google Workspace, and OpenAI begins rolling out text watermarking in the EU.

In partnership with

Built for the unstoppable.

The Dell Pro powered by Intel® Core™ Ultra with Intel vPro® adjusts power usage based on how you work, so your battery life never slows you down.

Today:

  • Mistral Large 4 scales open weights past one trillion parameters

  • EmbeddingGemma 2 brings multimodal search to on-device apps

  • Beam targets coding agents with 501B parameters and 23B active

  • OpenAI starts rolling out invisible text watermarking under EU rules

  • Claude can now edit Google Docs, Sheets and Slides directly

Open-Weight AI Gets Bigger, More Multimodal and More Efficient

Three new releases are pushing open and locally deployable AI in different directions. Mistral is scaling its flagship past one trillion parameters, Google is shrinking multimodal retrieval into a 740-million-parameter model for consumer hardware, and Reflection is using a sparse architecture to chase agentic performance with only a fraction of its total parameters active at once.

The common theme is efficiency rather than size alone. These systems are trying to deliver stronger coding, retrieval and multimodal capability without requiring every parameter to run for every token or every workload to leave the device.

DeepSWE 1.1 benchmark chart showing Mistral Large 4 Preview scoring 62, ahead of Beam and Qwen3.8 Max but below the top Kimi result.

Mistral Large 4 is Mistral’s largest model yet, with roughly 1.05 trillion total parameters, 49 billion active parameters and a 1.6-billion-parameter vision encoder. The model accepts text and images, supports a one-million-token context window, and is designed for coding, agents, document work and multimodal reasoning.

Mistral reports 61.7% on DeepSWE v1.1, 59.4% on SWE Atlas Codebase Q&A and 28.3% on Terminal-Bench 4.0, producing a combined Coding Agent Index score of 49.8%. In a blind coding-quality evaluation, professional reviewers gave the preview a 3.74 average score, placing it behind Claude Opus 5 at 4.22 but ahead of Kimi K3, GLM-5.3 and GLM-5.2 in that test.

The public preview currently lists discounted API pricing of $0.68 per million input tokens, $0.07 for cached input and $2.09 per million output tokens. Mistral’s model page shows those rates at 50% below the displayed list prices, without announcing an end date for the preview discount.

Mistral Large 4 is available now through the Mistral Studio preview API. The downloadable weights are scheduled for release by the end of October, so the model is described as open-weight but is still API-only during the initial preview period.

Massive Text Embedding Benchmark chart comparing embedding models by size and code-task score, with EmbeddingGemma 2 scoring around 76 at roughly 300M parameters.

EmbeddingGemma 2 expands Google’s lightweight embedding model from text into a shared representation for text, code, images, video and audio. It is built on the Gemma 4 architecture, has 740 million parameters and is released under the Apache 2.0 license for local and commercial use.

The model has an 8,000-token context window and can process extended media including audio and video clips up to about 5.5 minutes. Its design is modular, so developers can use the text model alone or add the vision and audio encoders when their application needs those modalities.

Google reports a 61.36 multilingual MTEB score and 78.68 on the MTEB code benchmark, up from 68.76 for the first EmbeddingGemma on code. It also reports 64.64 on MIEB Lite for images, 50.67 on the MMEB video benchmark and 69.54 on MSEB retrieval for audio.

The full embedding is 768 dimensions, but developers can truncate it to 512, 256 or 128 dimensions to reduce storage and compute. Google says quality remains relatively strong at 256 dimensions, and the model is available through Hugging Face, Kaggle and Google AI Edge.

Benchmark comparison for Reflection’s Beam model, showing results across DeepSWE, Terminal Bench, SWE Bench, HLE, and CritPT against other open-weight models.

Beam is Reflection’s first open-weight model, built as a sparse mixture-of-experts system with 501 billion total parameters and 23 billion active for each token. Reflection trained it for coding, reasoning and agentic work, aiming to get frontier-scale capacity without paying the inference cost of activating the full model on every step.

The pretraining run used 23.8 trillion curated tokens from web and licensed datasets. Reflection says its reinforcement-learning phase generated more than 100 million rollouts on 10,500 Nvidia GB300 GPUs over four weeks.

Reflection reports 80.9% on SWE-bench Verified, 77.2% on SWE-bench Pro v2-Hard and 80.1% on Terminal-Bench v2.1. The results are more mixed on long-horizon coding: Beam scores 44.4% on DeepSWE v1.1, close to GLM-5.2 but well behind newer models such as DeepSeek V4.1 Flash, Kimi K3 and GLM-5.3 in Reflection’s own comparison.

Beam is still undergoing final red-teaming and evaluation. Early API access is available through a waitlist, while Reflection says the weights, technical report, model card and developer artifacts will be released later this month; API pricing has not yet been announced.

🧠RESEARCH

Kandinsky 6.0 Video introduces 3B Lite and 29B Pro diffusion models that generate five-second video with synchronized 44 kHz audio and lip-sync. A dual-stream CrossDiT links audio and video through bidirectional attention. The authors report stronger speech quality after reinforcement learning and release weights, code, and Diffusers integration for researchers.

RealCompanion benchmarks whether AI companions understand people across months of real conversations. The dataset contains 27,218 messages from ten relationships lasting up to 120 days. Results show relevant memories are usually distant and hard to identify, while three agent systems reconstruct personas at similar accuracy despite a 31-fold cost difference.

ProWAM improves robot control by predicting actions alongside ordered visual sub-goals instead of generating dense future videos. The model caches sparse planning features and replans with lightweight action denoising. It reports 85.8% on LIBERO-Plus, 75.7% on randomized RoboTwin, and 70% success in zero-shot real-world experiments across novel scenes and tasks.

📲SOCIAL MEDIA

🗞️MORE NEWS

OpenAI is introducing textGrain, an invisible statistical watermark that changes model word choices so eligible ChatGPT and Codex output can be identified by a detector. API customers worldwide can opt in now, while eligible ChatGPT and Codex text in the European Union will gain watermarking over the coming weeks; OpenAI says detection reached about 80% for 200-token passages and 95% for 400-token passages at a 1% false-positive target in one evaluation. The signal is not proof of authorship or accuracy and can weaken sharply after editing, so detector access is initially limited to approved researchers and expert organizations.

Claude for Google Workspace is now in public beta across paid Claude plans, adding a sidebar that can read selected content and edit the Google file already open in the browser. Users can require approval for each change or allow Claude to apply edits continuously, while separate Docs, Sheets and Slides connectors let Claude create and modify Google files from the Claude interface itself. Existing Claude connectors, skills and enterprise controls carry into the add-on, although the beta still has limits such as no Docs comment handling, no Connected Sheets editing and no linked-chart refresh in Slides.

OpenAI has released a new collection of mathematical results produced by an internal frontier model and published the work in a public repository with revision and citation protocols. The release includes formalizations of many proofs in Lean, ten summaries of the model’s reasoning, compute estimates and statistics on attempted problems, with OpenAI saying the average successful result used compute roughly equivalent to three hours of ChatGPT Pro thinking. OpenAI says it developed the release process with input from the independent mathematics advisory group hosted at the Institute for Advanced Study and plans to support workshops and programs focused on understanding major AI-generated results.

Independent limit testing found Anthropic subscriptions delivered roughly five times the API-equivalent value of comparable OpenAI plans when comparing Opus 5.5 with GPT-6.1 Sol under the tested agentic workload. The analysis also found OpenAI had cut the token allowance of its newly purchased $200 plan by about half while introducing a $500 tier, with older $200 subscriptions retaining their previous limits until October 29. The result depends heavily on the model, token mix and workload, so API-equivalent value is a comparison metric rather than a guarantee of how much useful work every subscriber will receive.

What'd you think of today's edition?

Login or Subscribe to participate in polls.

Reply

or to participate.