- NATURAL 20
- Posts
- AI Moves Down the Stack and Into Daily Work
AI Moves Down the Stack and Into Daily Work
PLUS: Gemini adds polished voice dictation on Mac, while Anthropic funds independent AI wellbeing tests.


Understand the real change behind your code
AI writes more code than ever. Reviewing it shouldn’t mean going through files in alphabetical order.
CodeRabbit Change Stack turns any PR into a layered walkthrough - overview page, timeline view, semantic diff and agentic chat - so you understand the real change behind your code.
Free in early access. Review your next PR with Change Stack today.
Review your next PR with CodeRabbit Change Stack Today
Today:
OpenAI: Jalapeño delivers faster, more power-efficient AI inference
Anthropic: Claude connects memory across Chat and Cowork
OpenAI: Admin plugin turns workspace requests into controlled actions
Google: Gemini adds intelligent dictation across macOS
Anthropic: $5 million program funds independent AI wellbeing tests
AI Moves Down the Stack and Into Daily Work
AI competition is spreading from chips to memory and administration, with companies trying to make systems faster, more continuous, and easier to control.
Today’s releases address three practical bottlenecks. OpenAI is attacking the cost and speed of inference with custom silicon, Anthropic is reducing the need to rebrief Claude, and OpenAI is bringing routine workspace administration into a conversation.
AI products are no longer judged only by what a model can answer; they are also judged by how efficiently they run, how well they preserve useful context, and whether organizations can govern them safely.
OpenAI published the first performance results for Jalapeño, its custom inference chip. Inference is the work a trained model does when it generates an answer, so faster and more efficient inference can reduce delays and operating costs.
On three public models, GPT-OSS 120B, DeepSeek R1 670B, and Kimi K2.5 1T, OpenAI reports 1.5 to 1.9 times more work per watt at peak throughput and 1.7 to 3.6 times lower end-to-end latency than the comparison systems. For highly interactive workloads, the reported performance advantage reached 2.1 to 4.1 times.
The company says Jalapeño is rated at 700 watts but stayed at or below 550 watts in the tested workloads. OpenAI used the public InferenceX benchmark and compared systems at matched response speeds, a more useful view than chip-level throughput alone.

Anthropic is merging the memory used by Claude Chat and Claude Cowork. A project discussed in Chat can now carry its saved context into Cowork, where Claude can act on documents and tasks without the user repeating the background.
Claude will add memory topics while a conversation is still happening rather than waiting until the chat ends. Users can inspect, edit, or delete what Claude remembers, and the feature is enabled by default for Free, Pro, and Max users on web, desktop, and updated mobile apps.
Anthropic says Claude will not save certain highly sensitive details, including government IDs, Social Security numbers, criminal history, and immigration status. Other sensitive topics are excluded by default but can be enabled by the user, with notifications when they are saved.

OpenAI introduced an Admin plugin for ChatGPT Work and Codex. Workspace administrators can ask questions about adoption and credit use, manage members and groups, diagnose permissions, adjust limits, and approve or deny spending requests from one conversation.
The plugin can also automate recurring checks. OpenAI gives examples such as routing pending usage requests to Slack or Microsoft Teams and granting feature access automatically when a request meets predefined rules, while sending exceptions to a reviewer.
OpenAI says the plugin acts only within the administrator’s existing role and workspace policies. Broader changes can be reviewed before they are applied, and the system reports what changed after an action completes.
🧠RESEARCH
ReWorld separates immediate control from long-term memory so world models can revisit earlier places without storing every frame. Using a fixed cache and retrieved landmarks, it streams 704-by-1280 video and reportedly recalls a starting view after 64 seconds.
Prime Agent is an open-source harness that preserves histories, memories, skills, prompts, and subagents across long tasks. Its coding environment standardizes execution, recovery, verification, and resource tracking. The paper report raising ARC-AGI-3 RHAE Best-at-one from 30% to 95.5%, while matching or beating other harnesses across several coding benchmarks in testing.
In a logic-puzzle study, cheaper AI access prompted more use, but participants who requested assistance performed worse after the tool was removed. A statistical model linked greater reasoning with larger skill gains. The result suggests assisted performance can overstate learning when AI substitutes for practice, though other tasks need testing.
📲SOCIAL MEDIA
🗞️MORE NEWS
Gemini for macOS can now transcribe natural speech into polished text inside any desktop window. Google says it removes filler words, handles mid-sentence corrections, and inserts formatted text at the cursor.
Anthropic launched a $5 million grant program for independent research into AI’s effects on user wellbeing. Grantees will receive funding, model access, and technical support, while publishing their evaluation tools as open source; the company says researchers will work independently.
Stability AI raised $76 million in Series B financing, bringing its reported total funding to $232 million. The round gives the Stable Diffusion maker more capital after years of leadership and financial turbulence, but funding alone does not establish product traction or long-term stability.
Keenable emerged from stealth with $26 million in seed funding and says it has indexed more than 100 billion web documents for AI agents. The company reports production use at several unnamed AI labs and inference providers, so the scale claim and customer adoption remain only partly verifiable.
What'd you think of today's edition? |
1
2


Reply