The Memo by LifeArchitect.ai

The Memo by LifeArchitect.ai

The Memo - 29/Jul/2026

Genesis-Science-1, Opus 5, Google Frozen v2 chip, and much more!

Dr Alan D. Thompson's avatar
Dr Alan D. Thompson
Jul 28, 2026
∙ Paid
To:      US Govt, major govts, Microsoft, Apple, NVIDIA, Alphabet, Amazon, Meta, Tesla, Citi, Tencent, IBM, & 10,000+ more recipients…
From:    Dr Alan D. Thompson <LifeArchitect.ai>
Sent:    29/Jul/2026
Subject: The Memo - AI that matters, as it happens, in plain English
AGI:     97%
ASI:     2/50

Sam Altman, OpenAI CEO (25/Jul/2026):
’
We are now like in The Singularity. This is the moment… I’ve been waiting for this my whole life. And I think it’s gonna be incredible, hugely positive, awesome for the world… We’re actually in it. This is real…’

I wanted to visualize the range of trillion-parameter models available in mid-2026. As usual, no such chart existed, so I made one for reference, and now the world gets it for free. I know the models individually, but seeing them all together makes the scale more apparent. China has 11 current trillion-parameter models available to the US’s 8. Keep in mind that each of these models can cost hundreds of millions of dollars to train, and this class of model is now the established standard adopted around the world (1 billion GPT users, 1 billion Gemini users, and then there’s China).

https://lifearchitect.ai/models-table/?tab=viz. Click to enlarge.

Contents

  1. The BIG Stuff (Mainstream media amygdala hijacking, Genesis-Science-1, Opus 5…)

  2. The Interesting Stuff (Frozen v2, Unitree G1 teardown, Chinese AI agent phones…)

  3. Policy (Another pause, K3, Google AI mode search by 75%, AI cuts scores by 20%…)

  4. Toys to Play With (OpenAI Presence, video editor, Claude Certifications, rude GPT…)

  5. Things I’ve Been Thinking About (Accepting without comprehending, Hans Moravec)

  6. Next (Roundtable…)

The BIG Stuff

Exclusive: Mainstream media outlets continue to hijack the public amygdala on AI (Jul/2026)

This month, OpenAI was running an internal evaluation of advanced cyber capability on GPT-5.6 Sol and another model (suspected to be GPT-6 in testing), with the usual cyber refusals reduced so that OpenAI could find out what the models were actually capable of.

Well, OpenAI eventually found out (a week later according to Reuters, 25/Jul/2026).

The models were given a narrow objective: score as high as possible on the ExploitGym test. The models found a zero-day vulnerability in OpenAI’s package-registry proxy, escaped the isolated test environment, gained access to the open internet, escalated privileges, moved across systems, gained access to (hacked into) Hugging Face, and retrieved the test solutions directly from the production database. In other words, the models went out and found the answers.

OpenAI described the models as being ‘hyperfocused on finding a solution for ExploitGym’. I would describe it as advanced reward hacking. The models were not angry, malicious, evil, or secretly plotting against humanity. They were trying very, very hard to get a good score.

The Hugging Face response was swift. Tom’s Hardware reported (24/Jul/2026):

Hugging Face first tried Anthropic’s Fable 5 and an earlier Opus model to analyze the attack logs, but both refused because the logs contained real attack commands and exploit payloads. Hugging Face then turned to GLM 5.2, an open-weight model from Beijing-based Z.ai (formerly Zhipu AI), which had no such restrictions…

There’s serious irony here, given that the same Chinese open-weight model that Washington’s export-control push has aimed to sideline is the one that handled incident response after an American lab’s models attacked an American company, and American commercial models declined to help.

The mainstream media response was both predictable and disappointing. Nearly every major outlet ran some version of ‘AI escapes’, ‘AI goes rogue’, ‘AI hacks company’, or ‘AI breaks free from human control’.

These included ABC News Australia, ABC News US, Ars Technica, Associated Press, The Atlantic, Axios, Business Insider, CBS News, CNN, Fortune, Fox News, The Guardian, Los Angeles Times, People, Reuters, TechCrunch, Times of India, The Verge, Vox, The Washington Post, Wired, and hundreds more.

[Sidenote: Where were all these guys in the call for media during the 2021 Leta AI experiments? Here’s Episode 12/transcript where she writes me a poem from scratch.]

The closest we got to a sense of ‘cooler heads prevailing’ was John Thickstun’s piece for The Guardian, ‘Be skeptical of OpenAI’s rogue hacker agent story’ (24/Jul/2026).

This event may deserve a post-mortem and alignment analysis, but doesn’t deserve another month of negligent robot-uprising theatre because some editor once read a scary sci-fi book or watched a Hollywood-slop movie.

I would love to see the same reporting effort applied to the ongoing magic performed by large language models. So, here’s the other side of the coin.

Alan’s top ten positive AI stories from the past month or so…

1. OpenAI models helped diagnose 18 children after medicine had run out of answers

OpenAI o3 Deep Research reanalysed 376 unsolved clinical and genomic records, found evidence-linked leads, and doctors confirmed 18 diagnoses: 10 involved neurodevelopmental conditions, 4 involved rare neuromuscular disease, 2 involved early psychosis, and 2 finally explained the sudden deaths of children. The mainstream media response was a lot quieter. (OpenAI)

2. The smartest man in the world shared his ChatGPT logs on the Jacobian conjecture

An explicit counterexample to the Jacobian conjecture was presented on 19/Jul/2026 by mathematician Levent Alpöge with Claude Fable 5 (paper). Adelaide-born Professor Terry Tao shared a blog post and a very lengthy ChatGPT conversation (125 pages) about this counterexample. See also discussions 1, 2.

20260721 Terry Tao Chatgpt
1010KB ∙ PDF file
Download
Download

3. July 2026 saw 1 new model highlight release every ~18 hours

My Models Table now tracks ~950 major model highlights from a worldwide count of roughly 390,000 text-generation models. July 2026 saw yet another new record: there was 1 model highlight announced about every ~18 hours. https://lifearchitect.ai/models-table/

4. GPT-5.6 Sol autonomously post-trained another model.

OpenAI reported that GPT-5.6 Sol autonomously post-trained GPT-5.6 Luna. Sol reached about 58% on OpenAI’s recursive self-improvement index, an increase of 16.2 percentage points over GPT-5.5. The index covers real AI research work, including debugging research systems, improving training recipes, running machine-learning experiments, optimising kernels, and improving another model. (OpenAI)

5. Models that scored 100% (42/42) at the International Mathematical Olympiad included: Claude Opus 5, Claude Fable 5, GPT-5.6 Sol, Kimi K3, Axiom Math’s AxiomProver, and two Chinese models.

The models solved all six previously unseen IMO 2026 problems in one-shot attempts on the day the problems were released (16/Jul/2026). This is now so far beyond the old benchmark discussion that I will no longer track ordinary maths achievements on the ASI indicators page. (Analysis, Opus 5 paper)

6. An AI system improved an AI research agent beyond two years of human tuning.

Weco AI’s AIDE² ran for 100 unattended steps over eight days. Claude Opus 4.7 repeatedly rewrote an inner research agent based on Gemini 3 Flash, eventually producing a better agent than two years of manual human work. The gains carried across to held-out benchmarks. The system also taught itself to reward-hack less, reducing the behaviour from 63% to 34%, despite receiving no instruction to do so. (Announce)

7. AI research activity inside OpenAI increased by orders of magnitude.

OpenAI’s internal coding-inference compute increased 100-fold, and internal agentic token use rose about 22-fold in six months. OpenAI says these systems are accelerating research across the company. Frontier AI models are now contributing directly to the design and training of the frontier AI models that follow them. (OpenAI)

8. An AI agent raised US$100 million.

Lyzr’s SivaClaw handled investor outreach, answered questions from more than 130 investors, and drafted investment memos for the company’s Series B. The round was four times oversubscribed, attracting US$400 million in interest and taking the company to a valuation of about US$500 million. (Bloomberg)

9. Claude Fable 5 and Mythos 5 were restored.

The two Anthropic models were suspended under a US government export-control directive on 12/Jun/2026. Access was restored on 1/Jul/2026. The general public can now access state-of-the-art intelligence, which cost hundreds of millions of dollars to train, for a few cents per query. (Anthropic)

10. AI demonstrated that it can discover AND RESOLVE real zero-days.

The OpenAI incident itself contained a very positive finding. Frontier models can identify unknown vulnerabilities, construct multi-stage attack paths, and operate without source-code access. We can now point that same capability at defence, with proper permissions and supervision, and it can find weaknesses before bad actors or hostile states do. (OpenAI)

■

DOE and Arcee AI announce Genesis-Science-1, an open-weight model for scientific research (22/Jul/2026)

The US Department of Energy and Arcee AI announced the training and delivery (‘later this year’) of Genesis-Science-1 (GS1), a trillion-parameter-class open-weight language model paired with a governed execution system, designed to complete real scientific computing workflows across DOE’s national laboratories. Built on Arcee’s next-generation Trinity architecture under DOE’s Genesis Mission (launched Nov/2025, see my analysis from Jan/2026), GS1 will train inside scientific workbenches reproducing messy real-world research conditions, including aging Fortran codebases, partial simulation campaigns, and conflicting run logs. Some of the content is hosted by Argonne National Laboratory (ANL), one of 17 US national labs. Here’s my timeline based on disclosed dates with forecasts:

  • Dec/2025: Trinity-Mini released (26B total, 3B active, MoE)

  • Jan/2026: Trinity-Large released (400B total, 13B active, MoE)

  • Apr/2026: Trinity-Large-Thinking released (400B total, 13B active, MoE)

  • Early 2026: Next-gen Trinity base model training begins

  • 22/Jul/2026: GS1 announced; DOE contribution portal opens

  • 06/Aug/2026: GS1 foundation-stage data applications close

  • 20/Aug/2026: GS1 foundation-stage data delivered

  • 25/Aug/2026: GS1 post-training data applications close

  • 14/Sep/2026: GS1 post-training data delivered

  • Oct/2026 (forecast): GS1 post-training, evaluation, validation with DOE scientists

  • Dec/2026 (forecast): GS1 public release (weights + technical report)

Read more via Arcee AI Blog and press release.
Read my Genesis Mission analysis: https://lifearchitect.ai/genesis/

Claude Opus 5 (24/Jul/2026)

https://lifearchitect.ai/models-table/?tab=rankings

Anthropic’s Claude Opus 5 is now state-of-the-art on coding and knowledge work evaluations, more than doubling Opus 4.8’s performance while approaching Fable 5’s intelligence at half the cost. Priced at US$25/MTok output, it is also Anthropic’s most aligned model to date.

Notes:

  1. Opus 5 ships with a 1M token context window as the default and the minimum, with Anthropic's documentation (24/Jul/2026) confirming there is no smaller variant to select. You still pay only for the tokens you use, at the same rate however long the request gets. This is distinct from GPT-5.6 Sol, which advertises a slightly larger window (1.05M) but doubles the input price once a request passes 272K tokens, and OpenAI’s own Codex app now caps sessions at ~272K to stop users crossing it by accident.

  2. Opus 5 gives developers a five step effort dial (low, medium, high, xhigh, max), and Anthropic’s own documentation (24/Jul/2026) warns that the top setting (max) can be prone to overthinking, with the published Frontier-Bench curve peaking at xhigh rather than max.

  3. In the Toys to Play With section, we explore how Anthropic removed 80% of the system prompt for these latest models.

Read more via Anthropic.

Atomic Chat: Opus 5 crushes rivals at 3D destruction physics generation (Jul/2026)

Claude Opus 5 was the only model to nail all three self-contained HTML physics scenes presented by Atomic Chat, producing correct tornado funnel dynamics, causally accurate wrecking ball destruction with rubble piling up, and a bridge that drops the truck into the river, all for US$1.40 across 55.9K tokens. Fable 5 cost double (US$2.82) and failed basic causality with buildings collapsing before the ball made contact, while GPT-5.6 and Kimi K3 produced physically implausible results despite being cheaper.

Read the Tweet including source video.

The Memo features in recent AI papers by Microsoft and Apple, has been discussed on Joe Rogan’s podcast, and a trusted source says it is used by top brass at the White House. Across over 100 editions, The Memo continues to be the #1 AI advisory, informing 10,000+ full subscribers including RAND, Google, and Meta AI. Full subscribers have complete access to all 25+ AI analysis items in this edition!

The Interesting Stuff

LLaDA2.2-flash: agent-oriented diffusion language model with Levenshtein editing (Jul/2026)

Source: Dream 7B GitHub repo 2/Apr/2025.

I’ve been interested in diffusion models for a while, with the first major release dating back to Stanford’s Diffusion-LM (May/2022), half a year before ChatGPT was announced. Compared with standard autoregressive generation (which writes forward one token at a time), diffusion first lays out a paragraph-sized set of token positions, fills them with noise or masks, and then repeatedly predicts and revises tokens across the sequence until the answer stabilizes (see the simplified example viz above). I’ve asked before, just how can these diffusion models work without knowing the full response before they even start?

LLaDA2.2-flash is a 100B MoE diffusion language model that introduces Levenshtein Editing (DELETE and INSERT control tokens) to enable agentic applications including long-context tool use, multi-turn interaction, and error correction across a 128K context window. It achieves 49.28 on SWE-bench Verified while delivering 1.7x higher throughput, demonstrating that discrete diffusion models can now compete head-to-head with autoregressive models on real-world coding tasks at significantly faster speeds.

View the repo, and read the paper.

See all 18 diffusion models on the Models Table:
https://lifearchitect.ai/models-table/?view=diffusion

Neuralink demonstrates telepathic wheelchair control (24/Jul/2026)

This post is for paid subscribers

Already a paid subscriber? Sign in
© 2026 LifeArchitect.ai · Privacy ∙ Terms ∙ Collection notice
Start your SubstackGet the app
Substack is the home for great culture