Guru99 AI Report โ€บ News Letter โ€บ Current Edition

OpenAI’s GPT-6 Astra Claims AGI, Benchmarks Disagree

ALSO: 71% oppose AI data centers, CrowdStrike IDs agents

Guru99 AI Report

Welcome to Guru99 AI Report!

Top Story: OpenAI’s president called GPT-6 Astra “AGI” โ€” but the benchmark behind the hype hides a major asterisk. This week we unpack what the demos actually showed, an AI catching hidden heart disease in two seconds, and Claude running its own safety research. Let’s dig in.

๐Ÿง  GPT-6 Astra: Impressive Agent, Contested “AGI” Claim

GPT-6 Astra: Impressive Agent, Contested "AGI" Claim
Brief Buzz: GPT-6 Astra can run your computer and handle long, complicated tasks, and the launch demos are astonishing โ€” OpenAI’s president floated the “AGI” label. But independent testers are far more cautious, and the benchmark that sparked the excitement carries a major asterisk.
  • In OpenAI’s launch demo, a single Astra session turned a yellow circle into a rocket and then a playable 3D game, and built a complete eBay listing from speech alone.
  • On OpenAI’s own hardest tests, Astra scored near 100% on ARC-AGI-3, about 98% on FrontierMath Tier 4, 100% on ExploitBench, and 72.6% on OSWorld 2.0 (versus Sol’s 65.7%).
  • The AGI buzz hinged on one number โ€” 99.9% on ARC-AGI-3 โ€” but ARC Prize also reported 62.7% for the same model, the gap coming down to OpenAI’s custom harness keeping hidden reasoning between steps. Head-to-head it’s 62.7% vs Sol’s 7.8%, not 99.9% โ€” still a record, but ARC Prize says it has no evidence of AGI.
  • On Artificial Analysis’s Intelligence Index, Astra only draws level with predecessor Sol and trails Fable 5.1, Opus 5, Fable 5, and Meta’s Muse Spark 1.3.
  • It’s efficient but pricey โ€” roughly a third of Sol’s tokens for coding, yet $10 / $50 per million input/output tokens (about 2.5x Sol’s price) โ€” and it’s OpenAI’s first model to hit the “critical” cyber threshold, finding two zero-days in testing.
  • Early access users were impressed but wary: Wharton’s Ethan Mollick said Astra worked “autonomously for days” on a source-based Library of Alexandria simulation โ€” a single anecdote, not a measured result โ€” and noted it drifts less on long tasks.
The 62.7%-vs-99.9% puzzle
The excitement surrounding the โ€œAGI eraโ€ was based on one figure: Astra achieved 99.9% on ARC-AGI-3 (compared to GPT-5.6 Solโ€™s 7.8%), a test involving the ability to learn entirely new puzzle games on the spot. However, ARC Prize, the organisation that runs the test, also provided a second figure for the same model: 62.7%. This difference is due to the harness โ€” the software placed around the model that controls its memory and tools. OpenAIโ€™s custom harness allowed Astra to keep its hidden reasoning between steps, whereas a neutral, shared harness did not. The model weights were the same, yet the scores were very different (and in the direct comparison it is 62.7% versus 7.8%, not 99.9%). Importantly, ARC Prize has not claimed to have achieved AGI; one of its co-founders stated that they do not yet have the evidence. Nevertheless, 62.7% is a record โ€” more than double the previous best. (Details about OpenAIโ€™s โ€œcriticalโ€ cyber designation and its staged rollout are also given in our launch story.)
What hands-on users actually found
Those early testers who had access were impressed but remained cautious. Ethan Mollick from Wharton stated that Astra carried out complex and significant work for him โ€œautonomously for daysโ€, creating a walkable, source-based simulation of the Library of Alexandria โ€” although the figure โ€œfor daysโ€ is based on a single anecdote and not on a measured result. A common point that came up was that Astra tends to drift less when working on longer tasks, staying with the ideas that are still current rather than bringing back earlier drafts.
๐Ÿ’ก Why Should You Care?
Astra shows off AI that doesn’t just chat but actually runs real software across long, multi-step jobs. For now it’s expensive and limited to a few users โ€” but with Claude Fable 5.1 out at the same price, there’s finally a clean head-to-head. The takeaway: judge these tools on your own work, not on the “AGI” label.

๐Ÿชช CrowdStrike Makes Every AI Agent Show ID

CrowdStrike Makes Every AI Agent Show ID
Brief Buzz: AI is now being used to run cyberattacks too fast for humans to stop. At its Fal.Con 2026 event, CrowdStrike showed off new tools to fight back โ€” the most important gives every AI agent a kind of ID badge, so access can’t be granted without one.
  • Every AI agent gets a unique ID that can’t be copied or faked, and is granted access only for the task it’s carrying out.
  • Because every action traces back to the person or system responsible, nothing happens anonymously.
  • A new AI-run system checks devices, logins, cloud services, and networks at once, cutting response time from hours to minutes (per CrowdStrike).
  • Other new software intercepts open-source code before it can run and spread inside a company.
  • CrowdStrike says its AI agents now outnumber its staff by 90 to 1.
๐Ÿ’ก Why Should You Care?
With attackers now moving faster than any human can respond, companies finally have a real shot at keeping up โ€” provided CrowdStrike’s tools work as advertised.

โšก America Turns Against AI Data Centers

America Turns Against AI Data Centers
Brief Buzz: AI data centers have become a political flashpoint. Roughly seven in ten Americans don’t want one built near them, and both parties are pushing back. New York and Texas have hit pause, and the backlash is starting to shape how AI gets built.
  • A Gallup survey found 71% of Americans oppose a data center near their home, with 48% strongly opposed.
  • In July, New York launched the nation’s first statewide moratorium, halting new hyperscale (50 MW and up) data centers for a year.
  • Texas then paused new approvals pending an audit โ€” though critics call it election-year theater.
  • Senators Sanders and Ocasio-Cortez introduced a federal bill to halt new AI data centers until safety rules are in place.
  • The EPA is now proposing to scrap the federal rule requiring public comment on data centers’ air-pollution permits.
๐Ÿ’ก Why Should You Care?
The data centers powering your AI tools strain local grids, water supplies, and utility bills. Whether they keep growing or stall will help shape the future of AI.

๐Ÿซ€ AI Catches Hidden Heart Disease in Two Seconds

AI Catches Hidden Heart Disease in Two Seconds
Brief Buzz: Researchers at Imperial College London have built an AI that reads a routine ECG in under two seconds and spots signs of heart failure and valve disease that doctors can’t see on the same trace. NHS trials are underway.
  • It was trained on 10.6 million ECGs, then tested on roughly 67,000 US patients across two groups.
  • It catches 81% of heart failure cases and up to 90% of valve disease cases.
  • ECGs are cheap and everywhere โ€” about a billion are done worldwide each year.
  • Testing is now running on 590 NHS patients across six hospitals in London and Bristol.
  • Routine NHS use is likely about two years away, pending approvals.
๐Ÿ’ก Why Should You Care?
A test you may already have had could flag heart disease early, speeding treatment and cutting the months-long waits for scans.

๐Ÿง  AI Assisted Live Brain Surgery, Saving a Man’s Sight

AI Assisted Live Brain Surgery, Saving a Man's Sight
Brief Buzz: In London, surgeons removed a brain tumor while an AI watched the live video feed, highlighting the nerves and blood vessels to avoid in real time. The patient kept his sight, and doctors are calling it a world first.
  • The AI analyzed the live surgical camera feed โ€” not the preoperative scans โ€” and flagged the areas to steer clear of.
  • Rhys Hibbert, 48 at the time, had an 11mm pituitary tumor found after his first seizure in December 2024.
  • Surgeons worked through his nose while the AI ran on a separate monitor as a guide.
  • The system was trained on hundreds of past operations โ€” more than most surgeons perform in a lifetime.
  • He could see again after about a week; the case is part of an ongoing clinical trial.
๐Ÿ’ก Why Should You Care?
For patients it means safer operations, with an expert “second set of eyes” catching dangers a surgeon might miss. The human stays in charge โ€” for now, the AI only advises.

๐Ÿค– Anthropic Had Claude Run Its Own Safety Research

Anthropic Had Claude Run Its Own Safety Research
Brief Buzz: A recent Anthropic study had Claude agents run safety research on their own. The AI examined 10 kinds of misbehavior โ€” from sycophancy to reward hacking โ€” and for each found fixes that improved safety without hurting the model’s other abilities.
  • Claude handled each flaw end to end โ€” literature search, method proposals, training, and testing.
  • The fixes closed between 26% and 96% of each “safety gap” (the distance from the current score to a perfect benchmark).
  • On deception, Claude cut the gap by 85% on average, versus 20% for six experienced human researchers.
  • The weaker Claude Sonnet 5 took 60 hours to improve a pre-release Opus 4.8, using about 15,000x less data than Anthropic typically uses.
  • Anthropic notes the human comparison isn’t fair โ€” the humans couldn’t iterate โ€” and only a narrow range of flaws was tested.
๐Ÿ’ก Why Should You Care?
Automating safety research could help oversight keep pace with AI’s progress โ€” but that will take real trust, grounded in results Anthropic itself is careful to qualify.
Krishna Rungta
Facebook LinkedIn Twitter/X

Hey! I’m Krishna Rungta

Founder of Guru99.com, Editor-in-chief & Technology Expert

Was this email forwarded to you? Sign up for free here.

Summarize this post with: