AI Takes On Longer, Harder Work

PLUS: ChatGPT connects Epic and nine health sources, and Runway previews an interface world model.

In partnership with

AI made PMs faster. Multiplayer mode is still broken.

A PM can summarize research, draft a PRD, and mock up a prototype before lunch. The hard part starts when the team has to decide what actually gets built.

Jira Product Discovery gives product teams one place to capture insights, prioritize ideas with consistent frameworks, and build living roadmaps stakeholders can rally around.

And because it’s connected to Jira, the context behind every decision stays with the work—so developers and their agents know not just what to build, but why.

AI helps PMs move faster. Jira Product Discovery helps the whole team build with confidence.

Today:

  • Claude Fable 5.1 cuts typical workload costs by 25%

  • Google’s Agentic video understanding cuts token use by up to 88%

  • Astra reaches “Critical” cyber capability with restricted access

  • ChatGPT connects Epic and nine healthcare sources

  • Solaris generates interactive visual interfaces in real time

AI Takes On Longer, Harder Work

Anthropic, Google, and OpenAI are pushing models into longer workflows, selective video analysis, and advanced cybersecurity, with tighter limits around cost and access.

Today’s releases focus less on chat and more on sustained work. Anthropic updated its coding and research model, Google gave Gemini a way to inspect selected moments in video, and OpenAI detailed controls around its strongest cyber model.

Benchmark table comparing Fable 5.1, Fable 5, Opus 5, and GPT-5.6 Sol across scientific research, coding, knowledge work, computer use, reasoning, and business workflows.

Anthropic released Claude Fable 5.1 and Claude Mythos 5.1. They use the same underlying model with different safeguards: Fable is generally available, while Mythos is limited to trusted-access programs for cybersecurity and life-science work.

Anthropic estimates that Fable 5.1 costs about 25% less than Fable 5 for typical token-billed workloads and up to about 45% less for highly agentic work. Cache reads now cost $0.25 per million tokens, 75% less; other rates remain $10 per million input tokens and $50 per million output tokens.

Anthropic reports 52.6% on Terminal-Bench-Science 0.1, compared with 24.7% for Fable 5, and 55.8% on Terminal-Bench 4.0. The scientific benchmark carries a standard error of plus or minus 3.5–4.5 percentage points.

Fable 5.1 can be used to discover software vulnerabilities but not to develop exploits. Anthropic reports that its newest cybersecurity safeguards block 60% fewer false positives than before.

Fable 5.1 is available across Claude platforms, Amazon Web Services, Google Cloud, Microsoft Azure, and the Claude API as claude-fable-5-1. Mythos 5.1 is initially limited to vetted cyberdefenders and life scientists at selected U.S. organizations.

Charts show agentic Gemini 3.7 Flash using fewer tokens and achieving higher accuracy across three video-understanding benchmarks.

Google launched agentic video understanding for Gemini 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite. It is available for uploaded and YouTube videos through the Gemini API in Google AI Studio and the Gemini Enterprise Agent Platform.

Instead of reading video at a fixed frame rate, Gemini can search and inspect selected segments across frames, audio, and transcripts. It can revisit fast action at different frame rates and retrieve short moments from long recordings.

Across Google’s video benchmarks, the feature reduced token use by up to 88%, cut analysis costs by up to 66%, and improved accuracy by up to 7%. Google notes that the largest gains appeared in long-form video.

Developers enable it by setting processing to “agentic.” It uses standard Gemini API token pricing with no additional feature fee. Google plans to bring it to the Gemini app soon and to Ask YouTube in the coming months.

ExploitBench chart showing Astra reaching 39% success at 76,000 output tokens, versus GPT-5.6 Sol reaching 11% at 138,000.

OpenAI now assesses Astra as meeting the “Critical” cybersecurity threshold in its Preparedness Framework, making it the first OpenAI model placed at that level.

Under the framework, that means a model can find unknown flaws and develop working exploits across many hardened systems with the right tools and access, without a person directing each step.

OpenAI reports that Astra scored 100% on ExploitBench. On an internal set of 20 recently disclosed high-severity V8 vulnerabilities, it achieved higher arbitrary-code-execution rates than GPT-5.6 Sol and found two previously unknown vulnerabilities used in an exploit chain. OpenAI is disclosing those flaws to maintainers.

Parts of Astra’s development and release were delayed while OpenAI strengthened safeguards. Advanced cyber capabilities will first go to a small group of testers, followed by access through Daybreak Blue for defensive work.

The deployment adds stronger refusals, misuse protections, and monitoring that can stop unauthorized activity. OpenAI also warns that these checks may slow, pause, or stop legitimate work. A system card is planned for launch.

🧠RESEARCH

BLOOM-WILT improves automated model audits by making rare behaviors easier to surface. It changes an auditor’s strategy across rounds and reweights likely outputs without retraining the target. Across four models and eight behaviors, it beat the baseline in 30 of 32 settings and changed earlier safety rankings.

This paper proposes agentic data cracking: structuring useful facts while an AI agent reads documents, so later questions can reuse them. On FanOutQA, an ideal structured store was 28 times cheaper. With one related follow-up per question, the method cut costs 53% while preserving accuracy in testing.

This study separates successful prompt injections into covert attacks, which leave no trace in the final answer, and overt attacks users might notice. Its ICoA method steers an agent back to the user’s task after an injection. Across four AgentDojo models, covert success rose 3.79–12.01 percentage points.

📲SOCIAL MEDIA

🗞️MORE NEWS

OpenAI added an Epic integration that brings authorized patient context into ChatGPT for Healthcare. It can summarize changes in a record and point clinicians back to supporting chart information; supported deployments can also place ChatGPT inside the EHR workflow.

A Healthcare Public Data plugin provides structured access to nine official sources, including ClinicalTrials.gov, CMS Coverage, RxNorm, DailyMed, and PubMed. In OpenAI’s evaluation, physicians rated 99.1% of responses safe across 4,363 ratings; those results are not an independent clinical validation.

Runway previewed Solaris, an Interface World Model that generates a visual interface frame by frame and responds to clicks, drags, and other user actions without first translating the design into code.

In Runway’s study of 250 participants and nearly 7,500 pairwise judgments, Solaris was preferred over a coded result in 61% of instruction-following comparisons and 71% of natural-behavior comparisons. Early access is available by request; Runway lists stable text, trust, long sessions, accessibility, and software integration as open problems.

Google’s September Android update will let Gemini save the location of untagged items, with an optional photo, in Find Hub. The feature is coming to Android 16 and newer devices in supported countries.

Guided vision uses the Gemini Live camera to describe surroundings and guide camera framing for blind and low-vision users. Google plans to bring it to Android 9 and newer devices.

AIR has raised $50 million across two seed rounds led by Sequoia and Greenoaks. Its platform discovers agents in company systems, continuously checks their tools and add-ons, and blocks interactions that fail security criteria.

AIR reports more than 20 customers and estimates that its platform filters out about 27% of the add-ons and skills it finds online. Those adoption and filtering figures come from the startup.

What'd you think of today's edition?

Login or Subscribe to participate in polls.

Reply

or to participate.