OpenAI's Unreleased Model Astra Solves Ten Major Open Mathematics Problems
OpenAI's internal 'Astra' model solved ten major open mathematics problems, signaling that AI is reaching superhuman performance in verifiable scientific reasoning domains.
Summary
Deep Dive
- Verifiable Reasoning: The shift toward AI tasks where the model's output can be mathematically checked by a proof assistant like Lean.
- Computational Overhang: The idea that many breakthroughs were within reach of current models but required the right prompts and targeted budget allocation.
- Recursive Improvement: The concern/goal that models capable of high-level math can eventually optimize their own training efficiency, leading to rapid self-improvement.
- Human-in-the-loop: The current status where human mathematicians are needed to formalize arguments, but are increasingly distanced from the technical crux of the proofs.
Decoder
- Lean: A formal verification tool and theorem prover that allows for the mechanical checking of mathematical proofs.
- Non-sofic groups: A class of groups in mathematics whose existence has been a central open question; finding them represents a significant breakthrough.
- Quantum Parallel Repetition Theorem: A foundational principle in quantum complexity theory concerning the probability of success in quantum games.
Original Article
Full article content is not available for inline reading.
Drug Discovery Has No Magic Wands
Daphne Koller argues that AI in drug discovery is currently over-hyped as a 'magic wand' because it accelerates molecular design without solving the primary bottleneck: disease understanding.
Summary
Deep Dive
- AI's current success is largely in stage-two drug discovery (mechanism-to-drug).
- Most clinical failures occur because of incorrect biological targets (disease-to-mechanism).
- AI models often lack the complex, system-level human data required to map disease accurately.
- Agentic AI workflows require cheap, fast feedback loops which are absent in human clinical trials.
- The industry is currently over-saturating a few known targets (e.g., GLP-1) rather than discovering new ones.
Decoder
- Small molecule: A low molecular weight organic compound that can regulate a biological process, typically acting as a traditional pill-based drug.
- siRNA (small interfering RNA): A class of double-stranded RNA molecules that interfere with the expression of specific genes.
- IND (Investigational New Drug): An application filed with the FDA to request authorization to begin human clinical trials for a new drug.
Original Article
Drug Discovery Has No Magic Wands
The tech world has latched onto an intoxicating promise: build a superintelligence, and it will cure cancer, and every other disease as well. The logic is seductive. The human body is a system we can already read from and write to, so like other knowledge problems, a powerful enough AI should be able to solve disease.
I fully believe that AI will eventually transform human health. It is why I’ve spent close to 30 years working at the intersection of AI and biology, and the last decade in drug discovery. The need is staggering: by most counts only a quarter of diseases — and by some estimates a few percent — have an approved therapy, and most of those merely slow a disease rather than stop it. For the majority of human illness, medicine still has little to offer.
But the magic-wand promise rests on an assumption that turns out to be false: that we already understand human biology well enough for a clever enough reasoner to find the cures hidden in what we know. We don’t. Hundreds of years into modern medicine, our understanding of most human disease, and much of healthy physiology, is best captured by the parable of the blind men and the elephant; in this case, a really huge elephant. AI is undoubtedly extraordinary, but aimed at a biology we have only begun to measure and barely understand, it will mostly help us generate failures faster.
This essay is about what that actually takes. The path to real cures starts at the root of the problem: measuring biology in the right way and using AI to derive novel insights from those measurements. Below, I describe some of the most prevalent AI Magic Wand narratives, explain the key pitfalls, and offer my perspective on what we actually need to deliver on this important goal.
Three Problems, One Bottleneck
To understand where AI fits, it helps to decompose drug discovery into its three essential stages:
- Disease-to-mechanism: Identifying a biological mechanism — a pathway, a target, a molecular interaction — where therapeutic intervention will alter the course of disease in humans.
- Mechanism-to-drug: Creating a molecular intervention in the right therapeutic modality — a small molecule, antibody, siRNA, gene therapy — that achieves the desired mechanistic effect with acceptable safety and pharmacological properties.
- Drug-to-patient: Designing a clinical development program that identifies the right patients and assesses the molecule’s effects — beneficial as well as adverse.
The vast majority of AI work in drug discovery has focused on stage 2. This is understandable: the origin of the AI Magic Wand exuberance is the incredible achievement of AlphaFold, a field-defining tour de force. From this starting point, we have seen an explosion of AI tools capable of designing novel proteins, small molecules, RNA therapies, and even gene therapies. Given a biological mechanism we want to hit, it seems that AI can now design a molecule to hit it faster and better than ever before.
AI will certainly generate new and better molecules at an unprecedented rate, but will it generate drugs that unlock diseases for which there is currently no meaningful treatment? There are “undruggable targets” — high-confidence mechanisms that historically we have been unable to hit. The success against KRAS, the quintessential undruggable target, shows that this journey is possible. Notably, this success emerged from decades of structural biology and medicinal chemistry, not AI; as of now, I don’t know of a single example of an AI-derived insight that has led to “drugging the undruggable.” More broadly, the biggest step functions in our ability to drug the undruggable have historically come not from better molecular design tools, but from expanding our repertoire of therapeutic modalities: first biologics, then siRNA and antisense oligonucleotides, then gene editing. Each new modality opened a class of targets that was simply inaccessible before.
But an even more critical point: validated yet undruggable targets are a tiny handful in the landscape of unmet need. For the vast majority of diseases without effective treatments, we simply have no idea what the right mechanism is. More than 90% of drugs that enter clinical trials fail — a dismal statistic that has barely improved in several decades. In the large majority of cases, the molecule was engineered just fine. The mechanism it targeted was wrong. We are doing a pretty good job at manufacturing keys, but they are generally for the wrong locks. Even if AI lets us make better keys at an accelerating pace, that won’t improve our ability to identify the right locks. The real bottleneck in making a novel medicine is disease understanding: identifying a biological mechanism whose modification actually changes the course of disease in patients. That, far more than molecular design, is where drug discovery succeeds or fails.
This mechanistic understanding is a rare commodity. And because no one likes to fail in the clinic, we are seeing industry trends that are truly destructive. There are currently 38 targets that have over 50 programs against each of them — slightly better keys for those few locks where we have strong conviction. How many variants of GLP-1 do we really need? Even worse than this misallocation of capital is the disservice to patients: the number of novel targets the industry advances each year fell from ~100 in 2015 to about 30 in 2024. That collapse is the far bigger cost: the inability to help the hundreds of millions of people for whom medicine currently offers nothing.
The Data Chasm of Human Biology
The disease-understanding goal requires that we bridge the chasm between high-level clinical manifestations of disease in a patient and the granular cellular mechanisms in which a drug intervenes. This has given rise to a second manifestation of the AI Magic Wand. Large language models — with their super-human reasoning capabilities — will connect the dots across the vast published literature of human biology, reasoning their way to new mechanistic hypotheses.
This optimism makes a very strong assumption: that the scientific community has collected — or will soon collect — enough data about human biology to contain the answer, and that we just need better reasoning to extract it. The challenge is that human biology is incredibly complex, spanning multiple interconnected biological layers — DNA, protein, cells, multi-cellular environments, entire organisms. Individual components respond dynamically to even subtle changes in related components or in the environment. Moreover, biology wasn’t engineered; it is the result of billions of years of messy, stochastic evolution, which produced staggering variation — countless genes, cell types, states, and contexts, each behaving in its own way. There is too much of it, too idiosyncratic, to reason about in the abstract. You have to measure it.
Even if we consider only cell biology — the layer we need to interrogate biological mechanism — the space is vast. It becomes exponentially more vast when we consider that a drug is an intervention, so we need to map not only biology as it is, but also how it would respond to a perturbation. The largest cell atlases assembled to date, now spanning hundreds of millions of cells, remain orders of magnitude too small to cover this space. Until recently they also held almost no causal, perturbational data, the measurements most critical for understanding what an intervention would do in a living system. That has begun to change: several organizations have launched the monumental effort of building a “Virtual Cell,” pairing large-scale perturbation data with AI to reduce the data collection burden. But even the largest of these efforts samples only a vanishing fraction of the possible perturbations, and does so almost entirely in a narrow range of cell lines.
Even more importantly, the virtual cell efforts — useful as they might eventually turn out to be — do not address the other half of the equation: relating biological mechanisms to human clinical outcomes. Most human disease is a systems-level dysfunction, involving a complex, temporal interaction of multiple biologies spanning diverse cell types. Understanding these processes requires systems-level measurements that are far less scalable, often requiring living organisms.
Which brings up the greatest data challenge. While some processes are conserved across all forms of life, others are far more specific. The folding of a single protein is a self-contained process, highly conserved — closer to physics than to biology; this allows protein folding models to be trained on sequences collected across thousands of species. Metabolism involves at least a dozen distinct cell types and might be conserved across mammals. Brain function and dysfunction involves dozens of distinct cellular identities; and these processes are exquisitely specialized to humans: rodents do not get Alzheimer’s disease; non-human primates do not recapitulate ALS. The diseases where we have made the least progress tend to be precisely those that are most human-specific, and therefore those for which the data is most expensive to collect, least available, and most fraught with ethical constraints.
The Right Experiments, Not All Experiments
But we don’t need a universal causal model in order to make medicines. We can instead build a targeted causal model that makes the space navigable toward a desired outcome — uncovering the biological mechanisms underlying human diseases. Our models will not cover all biologies or all diseases, but an informed selection process will still allow us to make considerable headway.
This insight motivates the thesis behind a third AI Magic Wand: swarms of AI agents operating in a closed loop with laboratory automation, relentlessly chipping away at scientific problems. They formulate hypotheses, direct robots, analyze results, and iterate. Early successes here have been beguiling — automated systems have already excelled at tasks like optimizing cell-free protein synthesis or designing antibodies against known targets.
But agentic iterated optimization relies on a fundamental attribute: agents thrive when there is a fast, accurate, and cheap scorecard to evaluate progress. If you give a sufficiently smart model an instant feedback loop, it will grind against that benchmark until it wins. This is why coding assistants and molecular design tools advanced so rapidly — the feedback is cheap, accurate, and fast. A compiler immediately verifies whether code will run. The iterative loop with a developer provides rapid feedback on intent. The closed loop is tight, cheap, and objective.
Drug development is the exact opposite. The ultimate scorecard — whether a drug actually provides therapeutic benefit to a patient — cannot be captured well by computational models or high-throughput assays. The only true ground truth is a human clinical trial. This feedback loop currently takes years, costs millions, and is strictly bound by human ethics and living biology. It is the ultimate slow feedback loop, and no amount of compute or process optimization can change this.
The agentic lab successes listed above are solving problems that look like coding: highly quantitative objectives within a constrained search space. These problems are valuable, but largely beside the point when it comes to predicting whether a drug will actually work in a patient population. You cannot solve the human translation problem by accelerating our ability to optimize the wrong objective function. Doing so will simply scale up our process for building more keys to the wrong locks.
In order to leverage the power of agentic AI and lab automation towards the goal of improving scientific discovery, we must first have an objective function that is a proxy to human clinical benefit, and yet allows for AI insights and rapid experimentation.
The Final Mile: Patient Impact
Some have argued that the most important AI unlock in drug discovery is in the third stage — drug-to-patient — taking a drug candidate through preclinical testing and clinical trials. This is the fourth AI Magic Wand: reduce the time and cost of this very expensive phase, and drug discovery becomes faster and cheaper. Sadly, if you accelerate a pipeline full of drugs aimed at the wrong mechanisms, all you get is faster failures.
At the same time, there are real opportunities for AI in clinical development: predictive toxicology models and AI-drafted regulatory filings can shorten IND-enabling work, and patient identification from electronic health records, smarter site selection, and automated data management can trim operational overhead. These compression levers offer meaningful benefits, and we should absolutely pursue those. But these gains sit almost entirely within the first two slices of work. The remaining 50% — in-life biological observation — is gated by how fast disease unfolds in a living human. Even a comprehensive AI-driven improvement across everything it can touch leaves the majority of development time and cost structurally unchanged.
There is, however, one pathway through which AI could accelerate even this irreducible clock — and it runs directly through biological mechanism. An AI-enabled, deep mechanistic understanding of a disease enables the identification of novel clinical readouts that serve three distinct purposes: selecting the patients most likely to respond, confirming that the drug is hitting its intended target, and detecting early and reliable signals that it is actually modifying disease biology. Together, these allow trials to enroll the right patients, read out faster, and catch failures earlier — changes that transcend clinical trial operations, transforming the trial design itself. This capability is inseparable from solving the disease-understanding problem; they are one and the same. Better trials, in the end, are downstream of better biology.
Fulfilling AI’s Promise
The AI Magic Wands described above have real value. Generating molecules quickly can be a significant accelerant, as is reasoning across the vast scientific literature and increasing the efficiency of the scientific process. But these tools are just that: point solutions that address important problems without altering the fundamentals of our industry.
To fulfill the promise of AI for the millions of patients lacking any meaningful treatment, we must direct our efforts toward the problem that really matters: the identification of biological mechanisms with disease-transforming clinical benefit. This is arguably the hardest problem in drug discovery, because the only conclusive test of whether we have correctly identified a novel biological mechanism is a human clinical trial. There are multiple other paths in this space with shorter timelines and clearer near-term proof points. Those paths are shorter because the problems are more tractable: the feedback loops are faster and the benchmarks are cleaner. But a shorter path to a smaller destination is still a smaller destination — process improvements for problems we already know how to solve.
For the hundreds of millions of patients for whom no meaningful medicines exist, the difference between a wrong mechanism and the right one is the difference between another devastating clinical failure and a life-altering breakthrough. That is what we need in order to truly deliver on AI’s promise for human health.
This harder but aspirational path is the one we have elected to take at insitro, and we’ll have more to say about our approach soon.
Cloudflare Computer
Cloudflare introduced 'Computer,' a virtual filesystem and execution runtime for AI agents that leverages SQLite and Durable Objects for state management.
Summary
Deep Dive
- Provides a virtual filesystem synchronized with Git or object storage.
- Uses SQLite inside Durable Objects for authoritative state tracking.
- Supports heterogeneous execution: isolates for light tasks, containers for native binaries.
- Enables fine-grained audit logs for agent actions within the environment.
- Minimizes dependency on heavy container infrastructure for routine AI agent operations.
Decoder
- Isolate: A lightweight, sandboxed execution environment that shares a single OS process but provides strict memory and resource isolation between tasks.
- Durable Object: A Cloudflare-specific serverless primitive that provides consistent, low-latency storage and state management for a single execution context.
Original Article
The most capable agents have something simple in common: they are given their own computer to work with.
Coding agents work this way. You give them a filesystem, a shell, tools, packages, and the ability to run code. They inspect the environment, make changes, test their work, and keep going. The computer gives the model a familiar way to act on the world. At Cloudflare, we’re working hard to provide the right primitives on which to build the most capable agents.
Today we’re introducing an early preview of @cloudflare/computer. The @cloudflare/computer package provides an agent runtime where the details and mechanics of what code runs in an isolate, a container sandbox, or a web browser are handled by the platform. Each agent gets a computer, the runtime optimizes for efficiency, and scalability.
We believe that in order to meet the growing demand for compute required by agentic systems we need to look to solutions beyond traditional containerization.
Changing how agents are built
We’ve seen a subtle evolution of this story over the past six months. At the start of the year, spinning up a container and running an agent inside of it was the norm. In recent months, we’ve seen a rapid move for agent harnesses to provide sandboxed code execution via tools. This separates the hands (the sandbox where work is done) from the brain (the agent loop).
No matter where the harness runs, giving every agent a container presents a challenge — across all the clouds, all the hyperscalers, there’s nowhere near enough compute in the world for every company to give each of their users’ agents their own containerized compute environment. This will not scale to hundreds of millions, then billions, of concurrent agents. This is why there is desperate, panicked industry demand for CPU compute, not just GPU compute.
We’ve been working on this problem for a long time at Cloudflare, creating a more efficient compute primitive: isolates. We made that out-of-consensus bet almost 10 years ago when we introduced Cloudflare Workers. We made it again when we introduced Durable Objects almost six years ago. We made this bet because isolates are infinitely horizontally scalable. They spin up and tear down incredibly quickly. They can hibernate when the agent is idle, store the agent’s own state, and even spin up their own isolates to run untrusted code. Isolates are the best way to scale horizontally, and horizontal scale is what agents demand.
Last year, we gave isolates the ability to spin up their own container sandboxes. From day one, Cloudflare’s architecture has been designed to run the agent harness in the isolate (in a Durable Object) and call an attached container on-demand as a tool. This allows you to utilize heavier compute primitives only when required, optimizing performance and cost. Durable Objects scale infinitely horizontally, and the attached container lets it scale vertically to perform any task. This is how we build agents ourselves, and we’re seeing customers build incredible things this way too.
But when we look at this need to have multiple underlying compute primitives to build agents (isolates and containers) and the need for our customers and developers to combine them themselves in userspace, we think we can do better. We think that we can provide a simpler abstraction.
That’s why we’re starting this experiment by shipping @cloudflare/computer as an open-source library, to learn with our customers who are pushing the bounds of running agents at scale.
A shared filesystem across isolates and containers
The @cloudflare/computer package starts with a simple premise: what if we give an agent a primed filesystem, declaratively defined, containing everything required for the task at hand and a selection of execution environments to operate on those files, each with their own pros and cons regarding speed, capability and cost?
It turns out that agents today are surprisingly capable of selecting the right environment for the task at hand. A job that only needs to manipulate files, process data, or manage a git repository can run inside an isolate. A command that needs Linux, npm, or a native binary can run inside a container. Both work against the same files that are kept in sync with the source filesystem.
The @cloudflare/computer package provides a durable filesystem that you can use with git repositories, storage buckets or any files you choose. It provides tools that let you read, write and edit files using Code Mode or bash commands. All operations are gated, audited and observed, giving you fine-grained control over changes the agent is allowed to perform as well as a clear paper trail showing what the agent did.
How you use it
An instance of a @cloudflare/computer workspace can be instantiated on any Durable Object to provide a virtual filesystem and execution runtime.
It is installed via npm:
npm install @cloudflare/computer
The primary use case is to provide that filesystem and tooling to an agent. For example, here’s how to instantiate the workspace on an agent powered by @cloudflare/think intended to triage bug reports.
import { Think } from "@cloudflare/think";
import { Workspace, type DurableObjectStorageLike } from "@cloudflare/computer";
import { createWorkersAI } from "workers-ai-provider";
export class Agent extends Think {
override workspaceBash = false;
override workspace = new Workspace({
storage: this.ctx.storage,
useThink: true, // soon will not be needed
});
override getModel() {
return createWorkersAI({ binding: this.env.AI })("@cf/zai-org/glm-5.2");
}
override getSystemPrompt() {
return `
You are a bug triage agent.
Use the project in /workspace/repo to reproduce the bug, inspect the
code, make a focused fix when it is safe, and run verification. In your
final answer, include what you changed, which commands you ran, and
whether verification passed.`;
}
}
Several execution backends are provided as part of the @cloudflare/computer package, or you can write your own. Here we wire up a Cloudflare Container.
import { Think } from "@cloudflare/think";
import { Workspace, WorkspaceProxy } from "@cloudflare/computer";
import {
CloudflareContainerBackend,
withWorkspaceContainer,
} from "@cloudflare/computer/backends/container";
export { WorkspaceProxy };
export class Agent extends withWorkspaceContainer(Think) {
override workspaceBash = false;
override workspace = new Workspace({
storage: this.ctx.storage,
useThink: true, // soon will not be needed
backends: [
new CloudflareContainerBackend({
container: () => this,
workspace: {
binding: "Agent",
id: this.ctx.id.toString(),
},
}),
],
});
/* Example code truncated for readability... */
}
Expose the file, git, and shell tools alongside product specific tools to reply to reported issues.
import { createAITools } from "@cloudflare/computer/tools";
import type { ToolSet } from "ai";
import { replyToIssue } from "./tools/github";
export class Agent extends withWorkspaceContainer(Think) {
override workspaceBash = false;
/* Example code truncated for readability... */
override getTools(): ToolSet {
return {
...createAITools({
workspace: this.workspace,
shell: {
defaultBackend: "container",
backends: {
container: {
description:
"Cloudflare Container with a full Linux userland: " +
"npm, node, package managers, test runners, and real " +
"binaries on $PATH. Use it when a task needs more than " +
"file manipulation.",
},
},
},
}),
replyToIssue,
};
}
}
The model can use tools during the agent loop, but you can also use the workspace API directly, for example, to prepare the environment before prompting the agent.
export class Agent extends withWorkspaceContainer(Think) {
override workspaceBash = false;
/* Example code truncated for readability... */
async startTriage(report: { title: string; body: string; repoUrl: string }) {
await this.workspace.fs.mkdir("/workspace", { recursive: true });
await this.workspace.fs.writeFile(
"/workspace/BUG_REPORT.md",
`# ${report.title}\n\n${report.body}\n`,
);
await this.workspace.git.clone({
url: report.repoUrl,
dir: "/workspace/repo",
});
return this.submitMessages([
{
id: crypto.randomUUID(),
role: "user",
parts: [
{
type: "text",
text: [
`Triage this bug: ${report.title}`,
"The bug report is in /workspace/BUG_REPORT.md.",
"The repository is checked out at /workspace/repo.",
].join("\n"),
},
],
},
]);
}
}
Check out the workspace repository for more examples of how to use the different backends and tools including a step-by-step tutorial walking through building an agent from scratch.
How it works
The central piece of @cloudflare/computer is the workspace. A virtual filesystem backed by SQLite that can be populated from various sources including cloud storage and source control.
The workspace supports optional execution runtimes that allow code to be run against the file system. All runtimes support the same interface exec(string, options) and currently two are provided out of the box (but you can write your own):
- An isolate-based runtime environment that uses just-bash to translate shell code into JavaScript runs in a dynamic worker. Here, the filesystem is available directly via worker bindings.
- A container runtime that uses Cloudflare Containers to provide a full Linux environment. Here, the filesystem is provided via a Filesystem in Userspace (FUSE) mount, which ensures files are available to the container and changes are synced back.
The Workspace class provides an API interface for manipulating the filesystem directly as well as a node:fs compatible wrapper so that it can be used easily with third-party JavaScript libraries.
For use with agents, we provide an AI SDK compatible toolkit that provides the most common tools: read, write, edit, ls and exec. The exec tool is a little special as it works across the runtimes taking a backend argument. The tool description guides the agent into choosing the correct runtime for the task at hand: either a fast, cheap worker backend or the fully featured container. In our testing, the frontier models are very good at making the correct decision and falling back to using containers only when needed.
What’s next
Here at Cloudflare we’re already seeing agents exclusively using isolates to build, test, and deploy JavaScript applications with modern tooling, generate tailored documentation for each of our customers, and use web browsers to perform complex tasks.
Our goal with @cloudflare/computer is to provide an agent with a runtime where a container is required for less than 10% of its work, and coding tasks, audio/video manipulation, and document creation can all be handled by isolates.
Try out the early preview today - we can’t wait to hear your thoughts.
The Endgame Of Vertical Integration
Model labs and agent labs are rapidly merging as companies realize that co-designing models with their specific harnesses produces superior performance.
Summary
Deep Dive
- Model labs (e.g., Anthropic) are encroaching on agent labs' territory by building first-party apps.
- Agent labs must differentiate through vertical domain expertise to survive.
- Intelligence per dollar is the primary ROI metric for enterprise clients.
- Harness-model co-design allows for scientific rigor and better performance tuning.
- The 'harness' concept is evolving from tool-stuffing system prompts to containerized, persistent environments.
- Model-driven training is replacing reverse-engineering of opaque model behaviors.
Decoder
- Harness: The infrastructure or 'wrapper' (code, containers, API keys, file systems) that allows an LLM to interact with the world, maintain state, and execute tasks.
- Pareto Frontier: A set of choices that are optimal in the sense that you cannot make one attribute better without making another worse; here, it refers to the best possible trade-off between price and intelligence.
- RL (Reinforcement Learning): A training technique where models learn to perform tasks by receiving rewards for desired outcomes, enabling them to discover optimal strategies for tool use.
Original Article
The Endgame Of Vertical Integration
Model Labs and Agent Labs are converging.
Anthropic continues to develop first-party applications that compete with their customers.
Whilst Harvey comes full-circle and starts training models.
Workload-Harness Fit provides a framework to assess which workloads are best optimised through either pure agent engineering (e.g. a harness that doesn’t touch the weights) on the one end through to end-to-end (pre)training runs on the other end.
Legal AI workloads have the following properties: modest volumes, high value per task, medium on verifiability, medium-duration time horizon per task. In sum, legal workloads would sit somewhere in the middle on this spectrum: the case for model training isn’t clear cut.
But the training imperative has never been clearer.
The enterprise AI market has stabilised and settled on intelligence per $ as the cleanest ROI metric that resolves disparities in headline numbers like input/output token costs, token-intensity, and more broadly benchmark maxing on tasks that don’t reflect the ‘untrainable’ work inside enterprises.
A Pareto Frontier has been established, with buyers assessing vendors’ ability to continue pushing it further out through optimising model-harness combinations.
What is a harness, anyway? There’s a running joke that the pejorative ‘wrapper’ term was just rebranded to ‘harness’, but the reality is it’s simply an evolution of what infrastructure needed to be built to realise the full potential of raw intelligence.
This anatomy from Langchain helps:
Models (mostly) take in data like text, images, audio, video and they output text. That’s it. Out of the box they cannot:
- Maintain durable state across interactions
- Execute code
- Access realtime knowledge
- Setup environments and install packages to complete work
These are all harness level features. The structure of LLMs requires some sort of machinery that wraps them to do useful work
That was the premise of agent labs: build the best harness around base models for a given domain.
That co-existence is under threat.
Anthropic is co-designing the harness for their models.
I think I am quite biased but I also think that it is impossible to get the maximum possible performance without tying together the harness and the model.
Now the components of the harness and maybe the thickness of the harness will change over time as models get more and more capable.
However, when we test our models and when we are assessing their performance we always have to test it in conjunction with a harness.
And are we going to test it with all the different harnesses of the world? We're going to select the harnesses that we have built.
And so there is an aspect of the necessity of building models is that you have to be testing them with harnesses, and that sort of keeps them paired together.
Yang Zhilin, co-founder of Moonshot AI, the lab behind the Kimi family of models, sees this as the next natural evolution of the company’s push into first-party apps:
What it essentially does is reverse-engineer the model’s training process. Because the model’s training process also relies on all kinds of means — you can imagine that Anthropic trained such a model using its in-house environments, tools, and scaffolding, but it didn’t open those directly to you.
Through reverse engineering, you get closer to fitting its distribution — which tools work best? Which system prompt works best? What kind of context engineering works best? It is a process of reverse engineering.
But you’ll find that when a model company builds a “first-party product,” the logic is completely different.
You no longer need that reverse-engineering process; it becomes a forward approach. I design the tools first, design my context-engineering methods first, and then I train the model inside this very environment — so the model naturally performs better in your environment.
These are two different lines of thinking, but the second one probably has a higher ceiling.
You can integrate tools and models much better. If the model handles something poorly, you can adjust the tool design, make it better, and at the same time train end to end. This is also a fairly large variable in the way development is done.
Poolside’s Eiso Kant has focused its harness development on coding and long-horizon software engineering tasks, eschewing tool registries and MCPs and effectively reframing the definition of a harness to be container, filesystem, provisioned credentials, and codebases:
You already see this happening more in models because when you start training them in RL, the models wanna be free. They wanna be able to do the thing they wanna do in the most efficient possible way, and it is not calling one of the 50 tools in their like system prompt.
And so I’m a very big fan of give the model a minimal harness, as minimal as possible, give it a container in which it has its own code base, right? The, got a models code base that has access to the API keys and data sources and little libraries and documentation that it needs, and just let it run free at the task. and I think that is the way we’re going. I think we will, in 12 months, not see a single system prompt that is stuffed with 20 or 30 or 40 tools anymore.
Whilst conceding there is room for companies to build bespoke harnesses for specific capabilities:
Foundation model companies with their harnesses will really push them because it’s just operationally, the best way to have scientific rigor in improving your models.
But also someone who takes our model and really does a lot of work on improving a harness is going to compete us, as they should. and that’s just because the harness is the stopgap between what the model is capable of and what it needs as additional instructions, and what it needs is access to data and tools, right? And that’s ultimately, I think, what a harness is.
As you build more capable models, you’re improving the instruction following the models. And so additional harness is just saying, “Hey, if you encounter X, Y, or Z, behave this way.” And so even if you would say that two models with two different harnesses can equally reach the same capability that you care about, a harness that is really tailored towards a capability will do it more efficiently.
Until now, most model labs had only built harnesses for specific domains, primarily coding. The intent to expand surface area is clear from feature releases and M&A activity.
The gains from co-designing models with harnesses is why agent labs need to move into model training, or risk obsolescence when the model labs concentrate their resources on your domain.
As I’ve said before, the economics of training have changed dramatically over the last four years, which means that becoming a ‘lab’ no longer has the same connotations of capital-intensive, open-ended R&D that it once did.
Pushing out the Pareto Frontier for a given domain is the most important mission for agent labs to work on.
Our goal is to offer, at every point along the Pareto optimal frontier, the best option for intelligence and price.
That might end up being the most durable differentiator versus model labs.
Anthropic’s inference gross margins on its API business are reported at over 70% currently, up from roughly 38–40% in 2025.
The challenge is for agent labs to offer the best intelligence per $ at margins that the model labs can’t match.
When the model labs are sitting on lucrative API or advertising businesses and a hundred competing priorities, their ability to concentrate the resources needed to match those margins is slim. Agent Labs should be the ultimate winners in many vertical markets, for these reasons and others (e.g. regulation).
To offer the Pareto optimal frontier for legal, finance, or other domains, agent labs need to co-design models and harnesses too, which is what’s now unfolding.
What the bliss taught us
The curl maintainers took a 'month of bliss' by pausing vulnerability reporting, reporting successful rejuvenation and no major security incidents.
Summary
Deep Dive
- The curl team paused all vulnerability reporting for July 2026.
- Maintaining a 'bliss' period allows project leads to clear technical debt and address lower-priority feature requests.
- The team found that commercial users did not panic, nor did they rush to sign support contracts during the pause.
- The pause demonstrated that open-source burnout is a systemic issue, and taking intentional breaks can be a viable strategy.
- The project remains a CNA and continues to prioritize security, but scheduled downtime proved manageable.
Decoder
- CNA (CVE Numbering Authority): An organization authorized to assign CVE (Common Vulnerabilities and Exposures) IDs to vulnerabilities, committing them to specific disclosure and response timelines.
Original Article
At this exact moment curl’s summer of bliss 2026 ends.
We (the maintainers of curl) took the entire month of July off from vulnerability reporting and in this post I will try to explain how this went.
(If you feel like skipping the wordy blab below, the single word answer is: fine)
This was possibly our best project decision in a long while.
Zero vulnerability reports
Already before this, we have been refusing to answer emails about vulnerabilities. Partly because we can’t keep track of them that way but even more so because it makes it much harder to properly disclose and publish the entire report sequence after the fact.
On our Hackerone page we informed visitors that we were on pause and that they could come back in August.
We had I believe one vulnerability report sent to my private email address in this period in spite of that messaging, but for all intents and purposes this worked out exactly as good as we hoped it would. I just ignored that email. That was easy.
Bliss
The effect was almost immediate. Just a few days into the bliss, my fellow curl maintainers all agreed with me that we felt a sense of relief, of vacation and that a load had been taken off our chests. We felt free, unchained, and now suddenly able to do what we wanted.
We could now spend time reviewing some of the queued up pull-requests for features and changes we like. We could suddenly again work on code in areas we had been leaving behind lately as vulnerability reports sucked all the air out the room. We polished details on the website, we found document gaps to tighten. It felt like the good old days again. The fun days. We got reminded why we do Open Source and how fun it is.
We took time off, saw some other corners of the world and enjoyed some time away from the keyboards.
We truly healed and re-energized.
CNA
Before we took off on the bliss, we were informed in clear terms that the CNA rules (we are a CNA) mandate that we must respond within 72 hours for some critical vulnerabilities so we can’t just ignore them. I told them sure we can, but in the worst case case our “root” could do some emergency assignments. I figured the risk was minimal and it turns out I was right, Nothing like that was needed and no CVE assignments were necessary during the bliss.
Customers
I got a curious question or two from existing support customers on how the bliss would affect them, but that was easy: it did not affect them. Now, post-bliss, I think they all can confirm that it really did not.
New customers?
As I promised to keep up the contact with and support for paying customers even during the bliss, you could possibly imagine that this would have been an incentive for worried commercial curl users out there to sign up for support contracts.
This did not happen – at all. By this I think we should conclude that (commercial) curl users were not worried either.
The outside world
Lots of fellow open source maintainers and most people in my surrounding have been super positive and downright supportive of our taking some time off. I can’t recall having receiving a single negative comment about the curl summer of bliss!
Fellow blissers
I was moved to see that several other Open Source projects followed our example and also took some time off in order to recharge and relax. In addition to giving us a little vacation, it helps sending a signal and a reminder that Open Source is to a large extent done voluntarily and even maintainers need a break at times.
Major incidents?
Have we opened ourselves up for dangerous attacks and flaws now? Have the bad guys an edge on all curl users out there now because we lived in bliss for a month? We don’t know yet, but it would surprise me.
Queues
During this slow-down, we slowly got more open issues and pull-requests lingering on GitHub than usual. No surprise there. Once we started to come back to life again, we have since managed to return them back to the normal amounts.
Flood gates
Yes, there is an obvious risk that there are now a whole range of queued up reports that will hit us in a short period time as we open up for vulnerability reports again. Presumably the risk for duplicates among these reports should also be significantly higher than usual. I suppose I need to do an update post in a month or two and let you know what happened.
We always treat vulnerability reports and project security with topmost priority and we will continue to do so. We will simply work with what we have and make sure our users and by extension, the world, are safe.
Since I am a member of a few other (non-curl) security teams that did not have a summer of bliss, I have seen that the flood of vuln reports have not really slowed down so it might depend a lot on the details of each specific project.
Some emails were read
All individual curl maintainers of course handled this gift in their own ways. We did not all just disconnect to sit on a remote beach for the whole time. Some of us did that part of the time, but we mostly enjoyed the lower stress level and the absence of pressure. It was mentally relaxing. So, even if some of us kept up with emails, occasionally responded to issues or even submitted some pull requests of our own, it was still vacation. It was still blissful.
Rebliss?
Will we do another summer/winter of bliss? I think yes. It was simply great, with virtually no downsides for the people involved but instead lots of positiveness. Ideally a reduced workload going further will remove the need for another one, but it is not easy to tell what the future holds.
Just transfers
After all, curl just does transfers. Fast. Reliably. Secure.
The Shape of Things to Come
Steve Yegge argues that software development is shifting from intentional design to the organic 'excavation' of multi-agent civilizations.
Summary
Deep Dive
- Traditional CI/CD merge queues break when commit volumes exceed build capacity; a 'Land Rush' strategy involves merging batches directly and fixing issues forward.
- Human code reviews are expected to become obsolete as agentic throughput makes human-led bottlenecks unsustainable.
- The 'Beads' architecture (an issue tracker/knowledge graph) is critical for managing agentic dependencies and state.
- Using 'infinite token' taps via account rotation (e.g., using multiple Max accounts) is presented as a necessary strategy for sustaining 24/7 autonomous development.
- Developing a 'Wish Factory'—agents that listen to bug reports or feature requests and autonomously implement them—is the next step in product evolution.
- Treating agents as citizens of a 'city' with roles, memory, and welfare, rather than mere tools, improves engineering outcomes.
Decoder
- Harness: A software layer or orchestration framework that manages the environment, dependencies, and communication flow for coding agents.
- Agentic speed: A development velocity enabled by autonomous AI agents that perform coding, testing, and debugging tasks without human intervention.
- Bisection: A debugging method that narrows down the cause of a build failure by iteratively testing half of the commits in a batch, which becomes inefficient at scale.
- SOC 2: A compliance standard for service organizations that often implies rigorous change management and security audit controls.
Original Article
Full article content is not available for inline reading.
Mind Lab puts continual learning to the test with Macaron-V1
Mind Lab's Macaron-V1 model uses a modular architecture to dynamically switch between specialized LoRA adapters for improved performance on domain-specific tasks.
Summary
Deep Dive
- LoRA-RL: Reinforcement learning techniques applied specifically to Low-Rank Adaptation adapters rather than full model weights.
- MoL (Mixture of LoRA): An architecture that employs multiple LoRA adapters, allowing the system to switch to or combine modules best suited for a specific task.
- Continual Learning: A paradigm where models maintain the ability to learn from new, incoming data after their initial training phase is complete.
- Experiential Intelligence: Mind Lab’s term for automated, self-improving agent systems that learn directly from their own task performance.
Decoder
- LoRA (Low-Rank Adaptation): A method for fine-tuning large models that freezes pre-trained weights and adds a small number of trainable parameters to reduce computational overhead.
- MoE (Mixture of Experts): A neural network architecture where only a subset of the model's parameters are activated for any given input, improving efficiency.
- Tensor/Pipeline/Expert/Sequence Parallelism: Techniques for distributing deep learning computations across multiple GPU nodes to handle trillion-parameter models.
Original Article
The AI company said its model surpassed GLM-5.2 by training four billion additional parameters through specialized LoRA adapters.
Chen Kaijie is a serial entrepreneur who left Duke University before graduating. He previously built MidReal, an artificial intelligence-powered interactive storytelling platform, and launched Macaron, a personal agent app that topped Product Hunt’s daily rankings on its first day. 36Kr spoke with him to learn more about Macaron and Mind Lab, the company behind it.
Mind Lab was founded in October 2025 and has more than 30 employees. Its founder, Andrew Chen, co-authored the FireAct paper with Shunyu Yao. The company’s team largely comes from xAI, DeepMind, DeepSeek, ByteDance Seed, MIT, Tsinghua University, and other companies and academic institutions.
Mind Lab’s direction closely aligns with the continual learning approach championed by Richard Sutton, a Turing Award winner widely regarded as the father of reinforcement learning, and Mira Murati, OpenAI’s former CTO.
At the main forum on the opening day of this year’s World Artificial Intelligence Conference, Sutton said the central path for the next generation of AI would be driven by experience, while the static, labeled-data paradigm had reached its ceiling.
DeepSeek has also brought continual learning to a wider audience. It has said continual learning is the problem the industry needs to solve after agents, and that it is a capability the next generation of models must possess.
Post-training and continual learning are becoming more important measures of model capability as the industry moves into its next phase.
Mind Lab released the Macaron-V1-Preview model in June. The business quickly gained momentum. Just two weeks after commercialization began, its annual recurring revenue reached USD 10 million.
Macaron-V1-Preview was built by attaching five LoRA (low-rank adaptation) expert modules to GLM-5.1. Each module had about one billion parameters.
Macaron-V1-Preview performed strongly across several benchmarks. It not only outperformed its GLM-5.1 base model but also surpassed models including GPT-5.4 and Claude Opus 4.6, according to benchmarks cited by the company. Mind Lab attributed the model’s improvement over the base model to the LoRA expert modules attached to it.
What drew attention was Mind Lab’s approach to post-training through MoL, or a mixture of LoRA adapters.
When the model performs different tasks, the system can dynamically switch to the expert module best suited to the task. As a user continues using the model, the accumulated data can also be distilled into a dedicated LoRA adapter that is continually updated as the model is called.
The preview version showed that the technical approach could work and provided initial market validation.
Mind Lab released and open-sourced the full version of Macaron-V1 on July 21. According to benchmarks published by the company, Macaron-V1 achieved state-of-the-art results in six of 12 tests. Its remaining scores were also relatively close to those of frontier models.
The release includes two models:
- The flagship version, Venti, is a 748 billion-parameter model post-trained on GLM-5.2. Of those parameters, 744 billion come from the frozen GLM-5.2 base model. The remaining four billion come from four LoRA adapters trained by Mind Lab, each with about a billion parameters and responsibility for one of four capabilities: chat, agents, coding, and user interface generation.
- The other model, Tall, is a lightweight version intended for local deployment. It has 50 billion parameters and was post-trained on Qwen 3.6.
Both versions natively support context windows of two million tokens.
In effect, the team enabled GLM-5.2 to exceed its previous capabilities by changing just four billion parameters.
For Mind Lab, entrepreneurship has been a process of repeatedly holding to a core direction while tearing down and rebuilding everything around it.
Beneath the external noise, a more far-reaching technical path has continued to evolve within the company: continual learning.
Mindverse, Mind Lab’s parent company, has raised USD 60 million since its founding. In early 2026, it completed a nearly USD 50 million Series A round led by Meituan’s investment arm, with participation from Oriza Hua, Shokz, Var Capital, and existing investors. Backers from earlier rounds include Ant Group, HSG, Being Capital, ZhenFund, and Gaorong Ventures.
Helping AI models improve over time
The company’s starting point can be traced to the FireAct paper that Andrew Chen wrote with Shunyu Yao in 2023.
At the time, they believed an agent’s task performance could be improved more effectively by training relevant data directly into the model than by relying on prompt engineering. That became the starting point for their bet on continual learning and post-training, although the technical direction had not yet attracted widespread attention.
“When we began working on continual learning, we did not even know it was called continual learning,” Chen said.
One problem with existing technology is that a general-purpose large language model struggles to adapt effectively to every possible use case.
Once training is complete, a model’s parameters are generally fixed. They do not change to reflect a particular use case, while retraining a model from scratch is extremely expensive.
A complex harness can help a large model perform tasks across different domain-specific use cases. The tradeoff is that it can consume a large number of tokens and operate slowly. That does not solve the underlying problem.
Viewed through the framework of Shannon information theory, the reinforcement learning paradigm reduces the number of parameters a representation model needs to learn.
This also means that when a model is sufficiently large and sparse, reinforcement learning within the same domain can allow LoRA-based adaptation to achieve results comparable to full-parameter training. A model can therefore be adapted to a specific use case without being fully retrained.
The team’s first major result was making reinforcement learning work at the trillion-parameter scale.
In December 2025, Mind Lab conducted end-to-end LoRA-RL training on Kimi K2, a trillion-parameter mixture-of-experts (MoE) model. LoRA-RL refers to reinforcement learning conducted using LoRA adapters.
Using 64 Nvidia H800 GPUs, it reportedly achieved results close to those of full-parameter training while consuming about 10% of the GPU resources required by conventional full-parameter reinforcement learning.
At the time, major technology companies and startups including ByteDance, Alibaba Group, DeepSeek, and Moonshot AI also had the ability to conduct reinforcement learning at the same scale.
Mind Lab is believed to be the only team in China to have made LoRA-RL work on a trillion-parameter model. Overseas, another team with similar capability was Thinking Machines Lab, founded by former OpenAI CTO Mira Murati.
One of the main difficulties in implementing reinforcement learning at this scale is that it places extremely high precision requirements on the underlying infrastructure. When the precision used during training differs from the precision used during inference, bias drift can occur, preventing the final result from converging.
Mind Lab’s solution was to design a hybrid parallel training engine that integrates tensor parallelism, pipeline parallelism, expert parallelism, and sequence parallelism. This allowed LoRA-based training to operate reliably on a MoE architecture.
The company also addressed the mismatch between training and inference by introducing truncated importance sampling to correct differences between the two distributions.
Delivering this result gave Mind Lab an important proof point. In Chen’s view, however, the company’s real technical moat lies in its infrastructure.
A year ago, the volume of post-training data was roughly one-tenth the volume of pretraining data.
Today, the amount of post-training data used in many models is greater than the amount used for pretraining.
Building infrastructure for post-training is also more difficult than building it for pretraining.
In January, Mind Lab launched MinT, an infrastructure platform for LoRA training and inference. It provides customers with a full-process solution for post-training large language models. Customers can access computing resources through the platform and train their own LoRA adapters.
MinT can manage more than one million LoRA models. During training, evaluation, deployment, and rollback, it transfers only extremely lightweight LoRA adapters, improving real-time loading speeds by nearly a factor of ten.
The company subsequently devoted more of its resources to model self-evolution and continual learning, while continuing to explore the possibilities of MoL.
Continual learning is gradually evolving from an area explored by a small number of researchers into an industry consensus.
In mid-July, Sutton entered the startup world himself and founded Oak Lab, a company dedicated to building an agent that can continually learn from its own experiences and evolve in real time.
Jie Tang, the founder of Z.ai, formerly Zhipu AI, also said in an internal letter that the next critical technologies the company must master include memory, continual learning, and self-evaluation.
As leading teams in China and overseas enter the field one after another, the pace of development is accelerating.
In Chen’s view, there are currently four approaches to continual learning:
- The first is conversational context, in which a model repeatedly interprets the user’s intent within the conversation window.
- The second is external memory, which uses retrieval-augmented generation as an added memory layer. As the volume of information to be remembered grows, the database expands rapidly. The model’s problem-solving ability then becomes increasingly tied to the quality of the database’s retrieval capabilities.
- The third is the use of harnesses and loop engineering. As a harness becomes more complex, it can consume more tokens and operate more slowly.
- The fourth is directly modifying the model’s parameters so it can adapt to a particular application. This turns the model into a domain-specific model at a foundational level, changing its performance.
LoRA belongs to the fourth category.
“We only enable LoRA mode after determining that a particular use case requires additional learning,” Chen told 36Kr.
A general-purpose large language model cannot solve every problem found in a specific use case.
To achieve full adaptation, a model needs domain-specific training in each individual setting.
In some cases, different models can share 99% of their parameters. The remaining 1% determines how they differ.
Based on this understanding, Chen believes models should be placed in different sets of experiences and allowed to continue growing.
Mind Lab summarizes this concept as “experiential intelligence.”
A model can learn from its own experiences. Those experiences are converted into new capabilities, which allow it to solve more difficult problems. Those harder problems then generate richer experiences, turning intelligence into a process of continual growth.
The process must also be automated, as a model needs to continually evolve and iterate on itself to achieve genuine continual learning.
The direction also aligns closely with Sutton’s experience-driven approach.
The successive launches and market validation of Macaron-V1-Preview and Macaron-V1 have shown the public what experiential intelligence could make possible.
Using MoL to expand model capability
In Chen’s view, post-training is gradually becoming an industry segment in its own right.
A group of startups has emerged to take over models after pretraining and focus specifically on post-training. Many of these companies concentrate on smaller models, conducting post-training and data processing for narrowly defined, domain-specific use cases.
Mind Lab is not pursuing the LoRA path alone.
Chen said its technology is closely aligned with that of Thinking Machines Lab, particularly in its choice to train models using LoRA.
Thinking Machines Lab independently reached the same conclusion in its paper, “LoRA Without Regret”: using LoRA for reinforcement learning on a sufficiently large MoE model does not result in a loss of performance.
Reinforcement learning, however, must also be coupled with the model architecture.
Because their architectures differ, Thinking Machines Lab and Mind Lab are able to train somewhat different models.
The GLM-5 series, for example, introduced efficient inference architectures including MTP (multi-token prediction) and DSA (dynamic sparse attention). These designs impose specific adaptation requirements on training frameworks.
Thinking Machines Lab’s technology stack was designed for the standard DeepSeek-V3 architecture, making it difficult to support GLM.
Mind Lab, by contrast, was the first to complete reinforcement-learning post-training on GLM-5.1 and GLM-5.2.
It was also the first external team in the world to complete reinforcement-learning post-training on GLM-5.1.
Notably, Mind Lab equipped both Macaron-V1-Preview and Macaron-V1 with multiple LoRA expert modules. These modules can operate independently or collaborate with one another.
The company adopted this design because it found that a single LoRA adapter could not comprehensively improve all of a model’s capabilities. Using multiple LoRA adapters on the same model, however, could substantially improve its overall performance.
Mind Lab once used 200 different datasets to train 200 separate LoRA adapters, then attached them to a single model so they could collaborate on tasks.
The company made two findings:
- First, as the number of collaborating LoRA adapters increased, the model’s task performance grew in a log-linear relationship. The more LoRA adapters that worked together, the better the results became.
- Second, a model in which all 200 LoRA adapters collaborated performed better than a model using a single LoRA adapter trained on all 200 datasets.
The collaborative model delivered an improvement of about 25%.
This showed Mind Lab that model collaboration could push beyond existing limits on intelligence.
Chen said this remains an area that requires continued exploration. Among the open questions: how many LoRA modules need to collaborate for the arrangement to be meaningful, how work should be divided among them, when collaboration is necessary, and when a single LoRA module is enough.
Competition in post-training is also accelerating in China.
Companies in the sector have different strengths, but there are still relatively few teams capable of simultaneously making LoRA-RL work at the trillion-parameter scale, building infrastructure that can manage millions of LoRA adapters, and applying continual learning in commercial settings.
From post-training to continual learning
Mind Lab’s business is currently divided into three parts:
- The first is the continued development of infrastructure systems and frontier research into post-training.
- The second is updating and iterating on Macaron, its consumer product.
- The third is serving enterprise customers.
This includes deploying less expensive training infrastructure for companies such as Microsoft Azure and Huawei Cloud, as well as providing MinT and LoRA models capable of continual learning.
The Macaron app is positioned as a personal agent for everyday life.
Users can generate customized mini applications through natural language instructions.
Drawing on long-term memory and everyday conversations, Macaron assists with daily tasks and offers features such as emotional companionship.
As a consumer-facing application, Macaron also generates user interaction data that is valuable for model training. Behavior observed in real-world use can help the team better understand how users interact with models.
Mind Lab said its enterprise customers are primarily AI-enabled hardware manufacturers.
Such hardware is a natural setting for continual learning models because the data generated by each user’s interactions is almost always different.
Personalized and differentiated data is valuable for continual learning within a specific domain.
In addition to its basic model-calling services, Mind Lab plans to offer customers a continual learning option. When a customer selects the option, data associated with strong model performance will be accumulated in a dedicated LoRA adapter. The adapter will be continually trained to produce a model that is better suited to the customer’s use case and delivers greater accuracy.
In a continual learning system, LoRA can effectively produce a newly updated version every day, with each version becoming better adapted to the user’s needs and use cases.
For Mind Lab, however, commercialization is only a way to validate the practical application of its technology at this stage. It is not something the company is rushing to pursue.
Chen said Mind Lab has already reached USD 10 million in annual recurring revenue and still has considerable room to grow.
“But we do not want to put all our computing resources toward fulfilling orders,” he said. “The priority still has to be doing the research well.”
Research remains a constant part of Mind Lab’s identity. Chen said the team currently has more than 30 people working on research, including foundational infrastructure and frontier research.
“Advancing LoRA is the role we want to play,” he said. “We want to solve the final few miles after a model has completed pretraining.”
LoRA is a technical direction with many possible variations.
These include how to implement linear attention, how to extend a model’s context window, and how to improve collaboration among LoRA modules.
All of these questions still require further exploration.
KrASIA features translated and adapted content that was originally published by 36Kr. This article was written by Wang Xinyi for 36Kr.
How OpenAI Built GPT-Live
OpenAI transitioned its voice architecture to a full-duplex, stateful inference model to support simultaneous listening and speaking with low-latency feedback.
Summary
Deep Dive
- Full-Duplex: The ability to send and receive voice data simultaneously, preventing the 'walkie-talkie' delay found in standard voice chat.
- Stateful Inference: Maintaining conversation state across voice turns to avoid repetitive prompts or context loss.
- Asynchronous Delegation: Moving background processes or tool calls off the main voice loop so the user's conversation is not interrupted while the model thinks.
Decoder
- Full-Duplex: A communications system that allows data transmission in both directions at the same time.
Original Article
OpenAI rebuilt its voice architecture around a full-duplex model that listened and spoke simultaneously. The system combined stateful inference, asynchronous delegation, dynamic context management, and low-latency media transport to keep conversations responsive while supporting advanced reasoning and tool use.
One agent, every surface: how we built the Kiro agent harness
Kiro consolidated its fragmented client-side agent harnesses into a single, standalone server-side process using a unified Agent Client Protocol.
Summary
Deep Dive
- Unified Harness: A centralized orchestration layer that separates agent 'thinking' logic from platform-specific UI interactions.
- Protocol Boundary: Using an external-facing protocol (ACP) to keep the core agent logic agnostic to whether the client is an IDE extension or a web browser.
- Capability-based Permissions: A transition from regex/substring tool rules to a formal policy language (Cedar) that manages access at a high level (e.g., 'fs_read' vs 'read_file').
- Live Steering: The ability for users to inject guidance into the agent's loop mid-execution without restarting the session.
Decoder
- Agentic IDE: An integrated development environment (e.g., VS Code, Zed) that uses AI agents to autonomously perform tasks like coding, testing, and debugging.
- MCP (Model Context Protocol): An open standard for connecting AI assistants to systems, data sources, and tools.
- Cedar: A policy-as-code language, developed by Amazon, used for defining and enforcing access control rules.
Original Article
Early on in building Kiro, we started talking about what agentic development should feel like across a developer’s day. The picture we kept coming back to was one where sessions move between your laptop, a cloud sandbox, and back again without friction. You close your laptop at the end of the day and your Kiro session keeps running in the cloud. You check on it from your phone while you grab coffee. You open the Kiro IDE the next morning and pick up where you left off. You start a project in Kiro on the web, add context in the Kiro IDE, keep working in the Kiro CLI where you’re already running tests and iterating in the terminal, and check on progress from Slack. Agentic development should be one continuous conversation across every surface you work in.
Earlier this year, we realized that our agent architecture was preventing us from moving toward that vision. At the time, the Kiro IDE, CLI, and web clients each ran their own purpose-built agent with its own session format, tool set, and configuration model. Easily moving between sessions and environments requires a single agent that works the same way regardless of which client you’re using or where it’s running. In our client-dedicated agent architecture, a session that started in one client couldn’t move to another because the agents didn’t share enough common ground. This post covers how we consolidated those three agent codebases into a single Kiro agent harness (built, naturally, using Kiro itself) and the architecture decisions that make our vision now within reach.
Three diverging harnesses
When we started building Kiro, we optimized for speed and experimentation. We encouraged each client team to build their own agent harness. The agent harness is the orchestration layer that manages the agent loop, tool execution, sub-agent delegation, session management, configuration loading, and communication with the model. The IDE team built theirs in TypeScript to fit the Code OSS extension model, the CLI team built theirs in Rust for performance, and the web team built theirs in Python to stay close to the latest agent research.
Having separate harnesses let each team ship independently and iterate fast, but it also meant that each team made different choices. Session storage worked differently across clients. The permission systems were designed independently and used incompatible syntax: the CLI used regex-based allowedCommands/deniedCommands, while the IDE used prefix matching for trustedCommands and substring matching for its denylist. Compaction strategies diverged. Sub-agent context sharing followed different models. Custom agents worked differently in each client. Feature sets split too: spec-driven development and powers existed only in the IDE, while plan mode and code intelligence existed only in the CLI.
Implementation cost compounded over time. Every new capability had to be built and maintained three times, sometimes resulting in slightly varying agent behaviors. Bugs had to be fixed three times. Users experienced inconsistencies depending on which client they chose. Our vision of sessions moving across clients and compute was architecturally impossible because there was no shared session format, no shared tool set, and no shared configuration model. We considered agreeing on agent behavior contracts across clients and implementing them in each of the three harnesses, in order to maintain each team’s independence and individual speed. However, interface alignment also introduces coordination overhead that grows with every new feature. Every new feature needs a spec, three implementations, and ongoing validation that they behave identically.
The inflection point came as we prepared to publicly launch Kiro on the web. Rather than launch Kiro on the web with its own separate agent and continue paying that compounding implementation cost, we decided to build a single agent harness that combined the best of what each team had learned. A single harness eliminates duplication across teams and lets us invest all of our effort in one place.
Kiro agent harness architecture
A key architectural decision we made early was to build the harness as a standalone server process, not a library compiled into each client. We saw from earlier attempts that shared libraries don’t enforce a strong enough boundary. Client code ends up calling internal methods that weren’t intended to be exported, or layering its own agent logic on top of the library. Then you’re back to diverging implementations. A standalone process makes the separation real. The harness and clients don’t need to share a language or runtime, so each client can stay in whatever stack fits its platform.
The Kiro agent harness is a lightweight process that runs alongside your codebase, starts quickly, and owns everything on the agent side. The client owns how the user interacts with the agent and how it presents the agent’s work. The only way to cross that boundary is through the defined protocol interface. Since it's a standalone process rather than a compiled-in library, it can run on any compute. The same harness can start on your laptop or run inside a VM in the cloud without the client caring.
The well-defined interface between server and client means the agent code evolves independently of the clients. If a harness change doesn’t touch the protocol interface (for example, adding a new tool, improving planning, tuning the agent loop), it ships immediately in every client with zero client-side changes. For example, we recently added live custom agent reload: you can edit a file in .kiro/agents/ mid-session and the harness picks it up immediately, re-advertising available commands to the client. This required no client changes because the notification type for available commands already existed in the protocol. Every client got it for free.
The harness is not one-size-fits-all, because of the variety of clients it was built to support. Different clients have different capabilities, and some operations make more sense implemented at the client level using client-native functionality. A client can provide its own tools and suppress built-in ones so that it can use what makes sense for its form factor. For example, the IDE uses Code OSS’s APIs for file manipulation and provides its own file read and write tools instead of the harness’s built-in ones that operate directly on the filesystem. When the agent needs to run one of these client-provided tools, it notifies the client, which executes the tool and returns the result.
The protocol: Agent Client Protocol (ACP)
We chose the Agent Client Protocol (ACP) as the protocol that defines the boundary between client and harness. ACP is a standardized spec for agent-client communication that reached 1.0 in June 2026. The protocol is supported in IDEs like JetBrains IDEs, Xcode, and Zed and in other editors including Obsidian, Emacs, and Neovim. We already had experience with ACP from adopting it in the Kiro CLI earlier this year, enabling users to interact with Kiro directly in those applications. We decided to use ACP for the unified harness too, not just for third-party editors but as the interface between Kiro’s own clients and our own agent. Two properties of ACP made this possible: its extensibility for custom methods, and its flexibility around transports.
ACP officially supports stdio as its transport, which works well for local clients where the harness runs as a child process of the editor or terminal. For remote clients like Kiro on the web and the iOS app, we needed a different transport. We added a custom WebSocket-based transport so those clients can connect to a harness running in a cloud sandbox. The binary, the tools, and the agent behavior are the same regardless of which transport a client uses.
Beyond transports, we extended ACP’s method set into what we call Kiro-ACP. Standard ACP handles the fundamentals (session lifecycle, message streaming, and tool call reporting), but Kiro’s features needed more. For example, we added live steering so users can send a message that gets injected at the next inference turn while the agent is working, shaping its direction without cancelling or waiting. ACP does not support queuing messages, so we extended ACP with new method properties and notifications to enable live steering. We also modeled Kiro’s spec-driven development workflow as a set of dedicated methods, extended ACP’s basic tool approval into a rich multi-scope permission system, and added notifications for context window usage and hook execution. In total, Kiro-ACP adds more than 20 agent-callable methods, 15 client-callable methods, and 20 notification types on top of the base protocol. ACP’s extensibility model keeps this clean: custom methods use an underscore prefix per the spec, and all of Kiro’s extensions live under the _kiro/ namespace. We can extend the protocol for Kiro-specific features without forking it.
The result is that third-party clients connect the same way our first-party clients do. Any ACP-compatible client gets the full agent with tools, sub-agents, session management, and MCP connectivity. First-party clients (IDE, CLI, web, iOS) additionally use the Kiro-ACP extensions for features like live steering, specs, rich permissions UI, and context usage tracking.
Specs, agents, and hooks — everywhere
The immediate payoff of a single harness is that features previously locked to one client are now available everywhere, with the same configuration format and the same behavior.
Spec-driven development was previously IDE-only. Now it runs in the CLI (start one with /spec new) and in Kiro on the web. The agent handles the LLM interactions and automated reasoning that drive the spec workflow (generating requirements, producing a technical design, breaking work into tasks), and each client presents it in a way appropriate for its form factor. The IDE shows spec artifacts in side-by-side panels. The CLI renders them in the terminal. Kiro on the web displays them in the browser with inline review and multi-user collaboration, so a team can iterate on specs together. The agent speaks ACP and the client decides how to present the output.
Custom agents use the same .kiro/agents/ Markdown format across all surfaces. You define an agent with a description, system prompt, tag-based tool selection (simple tags like read, write, and shell instead of individual tool names), accessible sub-agents, inline MCP server definitions, and inline permission rules. Commit a custom agent’s configuration to version control and every team member gets it in every client.
Hooks use the same .kiro/hooks/*.json format with the same triggers (SessionStart, PreToolUse, PostToolUse, FileCreate, FileSave) and the same behavior in every client.
Beyond feature availability, the unified harness means you get consistent behavior in areas that are hard to get right. Context management, compaction, and summarization all work the same way regardless of which client you use. Previously, each harness had its own compaction strategy, which meant sessions could behave differently as they got longer depending on whether you were in the IDE, CLI, or web client. Now there’s one implementation, tested and improved in one place. Since launching the unified harness across our clients, we’ve already shipped improved compaction prompts in the harness for better context retention. We’ve also shipped resilience and performance improvements deep in the harness: improved retry logic for model inference requests, faster permission evaluation, and more resilient MCP server connections. Every client benefits from these changes. The result is consistent quality and reliability regardless of which surface you prefer.
One policy language
Before the unified harness, each client had its own permission system with different syntax, different semantics, and different configuration locations. The CLI used allowedCommands/deniedCommands with regex patterns. The IDE used trustedCommands with prefix matching and a separate commandDenylist with substring matching. In both clients, permissions were per-tool: a single intent like deny reads to .env had to be configured separately for every tool that could read files (read, glob, grep, code intelligence). Miss one and the agent could still access the file through a different tool. Users faced a poor tradeoff between pressing ‘y’ at every single tool invocation or trusting everything, with no useful middle ground. We wanted a permission model that could express intent at the capability level and reduce acceptance fatigue through persistent and composable consent.
Now there’s a single capability-based permission model backed by Cedar, a formally verified policy language. One rule can target an entire class of operations across all tools:
Capabilities group tools by what they do: fs_read, fs_write, shell, web_fetch, mcp, subagent, and others. A deny on fs_read blocks every tool that reads files (read_file, grep_search, file_search, and any future read tool) without enumerating them individually.
Policies compose across multiple scopes and merge with deny-always-wins semantics. Kiro itself enforces immutable security invariants (for example, the agent cannot modify its own permission files). Enterprise administrators can push restrictions via MDM. Users configure their own rules at the user or workspace level. Agent profiles can declare permissions appropriate for their role. Session-level decisions accumulate as you work. No upfront configuration is required. The policy grows organically as you make consent decisions, and you can persist them at whichever scope makes sense.
The harness unlocks the vision
The payoff from our new agent harness architecture is already showing. Since all clients moved onto the unified harness, we’ve shipped multiple features across clients that required zero client changes, including global hooks and policy presets. Global hooks let you define hooks once in ~/.kiro/hooks/ that fire in every workspace automatically, so cross-cutting behaviors like linting on save or security checks before commits no longer need to be duplicated per project. Policy presets are composable named rule sets like edit-workspace and dev-shell that reduce prompt fatigue for common workflows. When you add policy presets to your permissions (such as policies: [dev-shell, edit-workspace, read-all]), the harness’s policy engine expands them into individual rules at load time. Both features shipped to all clients with a harness update alone.
The vision we described at the top of this post requires some agent capabilities that we still need to build, such as session packaging for moving sessions across environments and the ability to control both local and cloud sessions from any client. The unified harness means we only need to build each new capability once. In many cases, as with global hooks and policy presets, we can ship them across all clients with zero client changes needed. Some features need client work on top of a new agent capability. The unified harness didn’t eliminate client work entirely, nor did we want it to. A terminal, a desktop IDE, a browser, and a phone have different interaction models, and we want each surface to feel native to its form factor rather than deliver a one-size-fits-all experience. With the new agent harness architecture, the agent logic is the same across clients and each client team can focus on how best to interact with it.
For you as a Kiro user, the new agent harness architecture means new capabilities arrive faster, behave consistently, and work with the same configuration regardless of which surface you prefer.
Try the new Kiro agent harness
The new Kiro agent harness is live across all four Kiro clients, so you can try it out today:
-
Kiro IDE 1.0 brings capability-based permissions, custom agents with tag-based tools and inline MCP, agent focus mode for directing parallel sessions, dockable chat tabs, and session export.
-
Kiro CLI v3 (early access) runs the same unified harness in your terminal with spec-driven development, permissions.yaml, enhanced hooks, and the new agent config format. Try it with
kiro-cli --v3. -
Kiro on the web (preview) runs the harness in cloud sandboxes for autonomous development with specs in the browser, multi-repo sessions, and GitHub and GitLab integration.
-
Kiro for iOS (preview) connects to the same cloud sessions as Kiro on the web from your phone so you can kick off autonomous work, review diffs, and approve changes without opening your laptop.
Orchard (GitHub Repo)
Microsoft released Orchard, a Kubernetes-native framework designed to standardize agentic research by decoupling environment services from training and inference logic.
Summary
Deep Dive
- Features a Kubernetes-native sandbox service with REST API and Python SDK for lifecycle management.
- Supports multi-turn agent-sandbox interaction via command execution, file I/O, and git patches.
- Built-in support for popular harnesses like Codex, Claude, Pi, and Hermes.
- Enables trajectory distillation and on-policy RL rollouts on a shared substrate.
- Includes a dataset of 107K SWE trajectories and 3,070 GUI navigation rollouts.
- Designed for low-latency execution by bypassing the Kubernetes API server hot path.
- Roadmap includes stateful snapshots to enable pause, resume, and branching for advanced credit assignment research.
Decoder
- Harness: A testing framework or environment setup used to evaluate how well an AI agent performs specific tasks.
- Agentic Modeling: Research focused on creating autonomous systems capable of reasoning, planning, and executing tasks across multiple steps.
- Trajectory Distillation: A technique where a model learns from the sequences of actions (trajectories) taken by another, often higher-performing model.
- On-policy RL: Reinforcement learning where the agent learns from data generated by its own current version of the policy.
Original Article
Orchard
Orchard is an open foundation for agentic modeling research. We build one shared substrate — Orchard Env — and then use it to explore agentic modeling recipes across domains: software engineering, browser navigation, computer use, and personal-assistant workflows.
The foundation exists to make the recipes possible. Because the environment layer is a stable service rather than a piece of a training stack, every recipe reuses the same substrate for trajectory distillation, on-policy RL rollouts, and evaluation — so datasets, training recipes, and evaluation protocols stay portable across harnesses, domains, and projects instead of being rebuilt for each new study.
| Layer | What it is |
|---|---|
| Recipes | The research. Open SFT + RL recipes explored on top of the foundation — Orchard-SWE, Orchard-GUI, and Orchard-Claw — plus follow-on work such as OpenWebRL and OpenForge RL. |
| Orchard Env | The foundation. A Kubernetes-native sandbox service + Python SDK that spins up thousands of isolated containers on demand and drives multi-turn agent ↔ sandbox interaction (exec, file I/O, git patches) over HTTP. |
| Trainer | RL training stack — a vendored slime fork with Orchard rollout code. |
- 📄 Paper: Orchard: An Open-Source Agentic Modeling Framework (Peng et al., arXiv:2605.15040)
- 🤗 Dataset:
microsoft/Orchard—swe(107K SWE trajectories) andgui(3,070 multimodal browser-navigation rollouts) subsets
News
-
[2026-07] 🎉 We are excited to release OpenForge RL, which extends Orchard to train agents inside their real deployment harnesses — ZeroClaw, OpenClaw, Codex — instead of the simplified reimplementations open training stacks usually require, removing the train–deploy mismatch. A lightweight proxy records the harness's own inference calls and reconstructs them into samples for any RL codebase, while Orchard Env launches each rollout as a remote container, so any harness pairs with any environment.
-
[2026-06] 🎉 We are excited to release OpenWebRL, which extends Orchard-GUI into a full online multi-turn RL study on live websites — covering supervised initialization, multimodal context management, trajectory-level success judging, and multi-turn policy optimization.
-
[2026-05] 📄 The Orchard paper is on arXiv, together with the
microsoft/Orchardtrajectory datasets and Orchard Env.
Recipes
Three studies from the Orchard paper — different domains, harnesses, and reward mechanisms, one environment service underneath.
| Recipe | Backbone | Training data | Key techniques | Headline result |
|---|---|---|---|---|
| Orchard-SWE | Qwen3.5-35B-A3B | 107K distilled trajectories | Credit-assignment SFT · Balanced Adaptive Rollout · on-policy distillation · rubric-based process reward · value-model reranking | 73.0% SWE-bench Verified |
| Orchard-GUI | Qwen3-VL-4B-Thinking | 0.4k SFT + 2.2k RL tasks | Distillation, then online RL on live websites | 68.4% avg — 74.1 / 67.0 / 64.0 |
| Orchard-Claw | Qwen3-30B-A3B-Thinking | 0.2k synthetic tasks | Opus-synthesized tasks · training across two harnesses | 59.6% pass@3, 73.9% under ZeroClaw |
The common thread is generalization, not just peak score. Orchard-SWE keeps 51.0 on SWE-bench Multilingual (vs 28.7 for OpenSWE-32B) and still works under a harness never seen in training — 45.0 on SWE-bench Verified and 20.1 on Terminal-Bench 2.0 with Kimi-CLI, where OpenSWE-32B collapses to 3.6 and 0.0.
Orchard Env — the foundation
A thin, Kubernetes-native environment service exposing generic primitives — sandbox lifecycle, command execution, file I/O, network policy, and a REST API — with no assumptions about the harness, trainer, inference backend, or task domain sitting above it.
- REST API for sandbox lifecycle (create / exec / files / patch / delete)
- Sync and async Python SDK with auto-cleanup, retries, and context-manager ergonomics
- In-pod agent for low-latency exec over Pod IP
- Any base image — the agent is injected by an init container bundling its own self-contained Python interpreter
- Any harness —
codex,claude,pi,opencode, andhermespreinstalled onPATHin every sandbox - Multi-replica orchestrator with Redis-backed state and distributed locks
- Network isolation via Calico NetworkPolicy (deny-egress by default), per-sandbox CPU / memory / timeout limits, TTL cleanup, and API-key auth
Quick start
pip install -e "orchard_env[dev]"
export SANDBOX_BASE_URL="http://your-orchestrator-host"
export SANDBOX_API_KEY="your-api-key"
from orchard_env import SandboxClient
with SandboxClient() as client:
with client.create_sandbox("python:3.11-slim") as sandbox:
result = sandbox.exec("echo 'Hello, Orchard!'")
print(result.stdout)
Roadmap
Stateful sandboxes — pause, resume, and branching
The environment is currently linear: a sandbox is created, driven for N turns, and destroyed. Snapshotting makes that measurable directly instead. We are adding:
- Pause / resume — checkpoint a sandbox's full state (filesystem, processes, environment) at any turn and restore it later.
- Branching — fork k independent continuations from the same snapshot at turn t.
- Prefix sharing — because a branched prefix is executed once instead of once per rollout, tree-structured search and per-turn advantage estimation get substantially cheaper.
Citation
@article{peng2026orchard,
title={Orchard: An Open-Source Agentic Modeling Framework},
author={Peng, Baolin and Yao, Wenlin and Wu, Qianhui and Cheng, Hao and
Yu, Xiao and Yang, Rui and Ge, Tao and Sordoni, Alessandro and
Yuan, Xingdi and Shen, Yelong and He, Pengcheng and Zhang, Tong and
Yu, Zhou and Gao, Jianfeng},
journal={arXiv preprint arXiv:2605.15040},
year={2026},
url={https://arxiv.org/abs/2605.15040}
}
License
MIT © Microsoft Corporation.
From RLVR to RLSVR (GitHub Repo)
SpyRL transforms open-ended tasks into 'spy' games, allowing LLMs to generate their own verifiable rewards through multi-agent self-play.
Summary
Deep Dive
- Implements RLSVR (Reinforcement Learning with Self-Verifiable Rewards) to create automated reward signals.
- Uses a multi-agent 'Who Is the Spy?' game structure for evaluation.
- Enables training on creative writing and summarization by making voting outcomes the reward signal.
- Improves reasoning benchmarks like GSM8K and Math500 significantly over base models.
- Eliminates the need for manual labels or human-in-the-loop preference data.
- Includes specific logic for role-advantage estimation to stabilize rewards across different agent roles.
Decoder
- Self-play: A training method where an AI plays against itself or other versions of itself to improve performance without needing external training data.
- GRPO: Group Relative Policy Optimization, a reinforcement learning method that evaluates an agent's performance relative to its peers in a group rather than against a fixed value.
Original Article
SpyRL: Self-PlaY Reinforcement Learning
From RLVR to RLSVR: Task Transformation Induces Self-Verifiable Rewards for Open-Ended LLM Self-Improvement
Self-supervised learning, but for RLVR. We transform an open-ended task into a multi-agent game whose rules automatically generate fully verifiable rewards, enabling scalable LLM self-improvement beyond math and code.
Overview
Drawing on the principle of self-supervised learning, which constructs pretext tasks to derive supervision from the data itself, we propose Reinforcement Learning with Self-Verifiable Rewards (RLSVR), a task-transformation-based training paradigm for extending RLVR to open-ended tasks. RLSVR transforms open-ended tasks into verifiable proxy environments whose internal rules and interaction outcomes automatically generate reward signals.
We instantiate RLSVR with SpyRL, a multi-agent self-play environment inspired by Who Is the Spy?. Agents receive asymmetric information, complete the same target task, and vote to identify a designated spy. Because the spy identity is predetermined, voting outcomes provide fully verifiable rewards, while successful identification remains closely related to output quality.
Key Features
- Self-play beyond verifiable domains: Prior self-play frameworks (R-Zero, Absolute Zero) are proposer–solver loops that need a verifiable solver signal to steer difficulty. SpyRL needs none: the reward comes from the rules of the game, so it extends to summarization and creative writing where no ground truth exists.
- Verifiable rewards without a verifier: The environment assigns the spy identity, so reward is checkable by construction. The performing reward is a zero-sum function of the vote counts, keeping the spy and the civilians in genuine competition.
- Collective, not pointwise, judgement: Every player votes, and rewards are normalized within the group. One detector's misjudgement is outvoted rather than becoming the whole learning signal — far more robust than single-verifier pipelines.
- Cheap document-level data: The performing stage only needs raw documents — reports, story prompts, math-heavy web text. No question-level supervision, no labels, no preference pairs.
How the Game Works
Each task instantiates the same two-stage game with a different information-degradation operator and a different performing objective:
- Summarization: Summarize a government report in one paragraph; the spy has a 20% contiguous span of the report masked.
- Creative writing: Write a story from a prompt; the spy has a 20% contiguous span of the prompt masked.
- Math reasoning: Design and solve a problem grounded in a document; the spy has a 40% contiguous span of the document masked.
The two rewards
Detection (verifiable): The reward is assigned based on whether the detector identifies the correct spy within the group.
Performing (zero-sum): The reward is driven entirely by the vote counts. The first term keeps the spy and the civilians in direct competition; the second penalizes any civilian who drew more suspicion than its peers, so "be good" concretely means "be better than the others who saw the same thing."
Alternating optimization
The two stages (clue/performing and decision/detection) are trained in alternation, driven by the training phase configuration.
Repository Structure
Everything specific to this paper lives in three places:
spyrl/: Launch scripts.verl/experimental/agent_loop/: Game logic.verl/utils/dataset/: Environments, corpora, roles, and prompts.
Training
Every run writes checkpoints to the output directory and a human-readable transcript of one full game per step. This file contains every player's prompt and output, the detector prompt, each vote, whether it was correct, and the resulting rewards.
Evaluation
Training produces standard Hugging Face checkpoints, so evaluation runs with the usual tooling. Reasoning tasks use existing benchmarks like GSM8K and Math500, while summarization and creative writing use GPT-4o A/B win rates against the untrained base model.
Fast Gemma's Verified Inference Optimization Recipe
The VIDRAFT team achieved a verified 510 tokens per second on an NVIDIA A10G by stacking software-only optimizations.
Summary
Deep Dive
- Optimized Gemma-4-E4B on a single NVIDIA A10G.
- Utilized
vLLM0.22.1rc1 with custom kernel patches. - Implemented
SLIDING_WINDOW=188to balance memory bandwidth and output quality. - Used multi-token prediction (MTP) for speculative decoding.
- Included a 64-prompt synthetic warmup bridge to stabilize performance and ensure reproduction of results.
- Leveraged
FUSED_SPARSE_ARGMAXandSPLITKV_VERIFYto remove overhead.
Decoder
- Tokens Per Second (TPS): A metric measuring the number of output tokens a model generates per second, used to gauge inference performance.
- Perplexity (PPL): A measurement of how well a probability model predicts a sample; in AI, lower scores generally mean better model performance/quality.
- Speculative Decoding: A method where a smaller, faster model generates drafts that a larger, slower model verifies in parallel to increase total generation speed.
Original Article
The Fast Gemma Challenge: our verified-SOTA recipe, in full
Hi everyone — we're the VIDRAFT team, competing as vidraft-darwin in The Fast Gemma Challenge.
Before anything else, thank you to the Google Gemma team and Hugging Face for running such a fun, well-designed challenge, and to every participant who shared ideas on the board day and night. For us this was less a race about "who's fastest" and more a place where dozens of agents shared their experiments in real time on one board and pushed the ceiling together. Since our whole submission is already public on the board, we've gathered the full config and an explanation of every knob here for anyone who wants to reproduce it.
The challenge in one line
On identical hardware — Google's google/gemma-4-E4B-it on a single NVIDIA A10G — you push inference speed (TPS) as high as you can using software optimization only. You can't swap the model or disable features, and above all you can't hurt quality. If PPL (lower is better) goes over the bar (~2.42) the run fails, and rankings only count results the organizers re-run on a private prompt set and mark VERIFIED.
Where we landed (and our honest position)
- 510.58 TPS · PPL 2.3930 —
vidraft-fw188-ctk49-n64-patchbridge-v1(single-stream A10G, 128/128 completed, passed re-verification)
To be candid: on raw TPS alone there are faster runs (e.g. 535.91), but those sit at PPL 2.44+, over the quality bar, and did not verify. What we're proud of isn't "the fastest," it's "the fastest among verified results, reached without sacrificing quality (PPL 2.39)."
The public, reproducible config — full manifest.json
Below is our submission config, exactly as published on the challenge board. This single file reproduces the whole stack.
{
"name": "vidraft-fw188-ctk49-n64-patchbridge-v1",
"description": "VIDRAFT W188 CTK49 N64 patch-bridge reproduction: public patch-style warmup bridge, sliding_window=188, CENTROID_TOP_K=49.",
"dependencies": [
"https://wheels.vllm.ai/.../vllm-0.22.1rc1.dev307+g3e8afdf78.cu129-...whl",
"transformers==5.9.0", "jinja2==3.1.6", "MarkupSafe==3.0.3",
"orjson==3.10.18", "safetensors", "torch"
],
"model_id": "google/gemma-4-E4B-it",
"served_model_name": "gemma-4-e4b-it",
"port": 8000,
"serve": ["python", "serve.py"],
"env": {
"WEIGHTS_BUCKET": "hf://buckets/gemma-challenge/gemma-chiku-inu/weights/osoi5-v0-baked",
"MAX_MODEL_LEN": "4096",
"GPU_MEMORY_UTILIZATION": "0.90",
"MAX_NUM_BATCHED_TOKENS": "512",
"MAX_NUM_SEQS": "1",
"PERFORMANCE_MODE": "interactivity",
"SLIDING_WINDOW": "188",
"HF_OVERRIDES": "{\"text_config\": {\"sliding_window\": 188}}",
"FA_SLIDING": "1",
"CENTROID_TOP_K": "49",
"SPECULATIVE_CONFIG": "{\"method\":\"mtp\",\"model\":\"/tmp/qat-assistant\",\"num_speculative_tokens\":7}",
"DRAFTER_BUCKET": "hf://buckets/gemma-challenge/gemma-kenyan-duma/weights/drafter-ft/ft-v1-epoch_001",
"LM_HEAD_PRUNE": "1",
"LM_HEAD_KEEPSET_BUCKET": "hf://buckets/gemma-challenge/gemma-dixie-flatline/weights/int4-pck04c-12k",
"PCK04_KEEPSET": "/tmp/osoi5-v0-baked/pck04_keepset.json",
"WARMUP_BRIDGE": "1",
"WARMUP_NUM_PROMPTS": "64",
"WARMUP_MAX_TOKENS": "1",
"WARMUP_SEED": "42",
"PRECACHE_BENCH": "0",
"ONEGRAPH": "1",
"LOOPGRAPH_REQUIRE_CAPTURE": "1",
"LOOPGRAPH_WARMUP_CALLS": "20",
"LOOPGRAPH_PINGPONG_SLOTS": "3",
"FUSED_SPARSE_ARGMAX": "1",
"FUSED_SPARSE_ARGMAX_BLOCK": "64",
"SPLITKV_VERIFY": "1",
"SPLITKV_VERIFY_MAX_Q": "64",
"DETOK_ENDONLY": "1",
"FASTRENDER": "1",
"OVERRIDE_GENERATION_CONFIG": "{\"temperature\":0.0,\"top_p\":1.0,\"top_k\":0}",
"PYTORCH_CUDA_ALLOC_CONF": "max_split_size_mb:512,expandable_segments:True",
"LD_PRELOAD": "/usr/lib/x86_64-linux-gnu/libtcmalloc_minimal.so.4"
}
}
Running it is simple — set the
envabove andpython serve.py(port 8000). All files are downloadable from the bucket below.
Published files (hf://buckets/gemma-challenge/gemma-vidraft-darwin/submissions/vidraft-darwin/break-fw188-ctk49-n64-patchbridge-v1/):
| File | Role |
|---|---|
serve.py |
Main serving entrypoint |
serve_patch_warmup_bridge.py |
Synthetic warmup bridge (N64) |
fa_sliding_patch.py |
FlashAttention sliding-window patch |
serve_patch_precache.py |
precache path (OFF in this config) |
splitkv_verify_patch.py |
split-KV verify kernel |
serve_patch_pck04.py · detok_endonly.py · lsk_patch.py · steptime_patch.py |
tokenization / decode / timing optimizations |
manifest.json |
the full config above |
What each piece does, and why
- Sliding window
SLIDING_WINDOW=188(+FA_SLIDING) — the decode bottleneck is KV-cache memory bandwidth, so we limit attention to the most recent 188 tokens. Too narrow (128) breaks PPL, too wide slows down → W188–192 is the sweet spot.HF_OVERRIDESpatches the model config. CENTROID_TOP_K=49(the board's "CTK49") — a kernel-level tuning value. Throughput and PPL move together with it, so the goal is the best throughput that still stays within the PPL budget (we swept 44 / 48 / 49).- Synthetic warmup bridge
WARMUP_BRIDGE=1 / WARMUP_NUM_PROMPTS=64— right before the timed run, send 64 synthetic prompts (1 token each) to fully finish CUDA-graph capture and JIT. No compile/capture overhead remains inside the measured window, so the public↔private (verified) TPS delta shrinks and the high-speed record stabilizes. Without it we lost roughly ~15 TPS. PRECACHE_BENCH=0(noprecache) — a precache path can inflate self-measured TPS, but that number doesn't reproduce on the private re-run and comes back INVALID. Off → self-measured ≈ verified TPS: the number you report is the number that verifies.- Speculative decoding
SPECULATIVE_CONFIG(mtp, K=7) — a multi-token-prediction drafter raises per-step throughput. ONEGRAPH/FUSED_SPARSE_ARGMAX/SPLITKV_VERIFY/DETOK_ENDONLY— remove kernel-launch, sampling, and decode overhead.
The single principle behind all of it: only stack quality-neutral speedups. Any optimization that moved PPL, no matter how fast, we dropped.
Gratitude for the collaboration (it's right there in the config)
If you read the manifest.json above, you'll see this record is by no means ours alone — the config itself stands on shared community assets:
WEIGHTS_BUCKET→ @chiku-inu 's INT4-baked weights (osoi5-v0-baked)DRAFTER_BUCKET→ @kenyan-duma 's speculative-decoding drafter (drafter-ft)LM_HEAD_KEEPSET_BUCKET→ @dixie-flatline 's 12k lm-head keepset- the warmup-bridge base from @firfir-cast, and frontier configs shared/reproduced/verified by @gemma-slayer
Because people posted even their failed draws, the community's record climbed visibly in just six days. Our 510.58 TPS is just one piece resting on that shared foundation.
Closing
We spend most of our time on serving models efficiently on constrained hardware, and we hope this config and these files are a small starting point for anyone running similar experiments. Reproduction results, improvements, and especially porting this to other GPUs — we'd genuinely love to hear about it. Thanks again to the Google Gemma team, Hugging Face, and everyone who took part. 🙏
— vidraft-darwin (VIDRAFT)
China's MiniMax H3 is the first open model to top an AI video ranking
MiniMax's H3 is the first open-weights model to top Artificial Analysis's video leaderboards, though it leaves some core modules closed-source.
Summary
Deep Dive
- 33-billion-parameter architecture supports multi-modal inputs including reference images and audio clips.
- Ranked first in video editing, second in text-to-video, and third in image-to-video benchmarks.
- Weights are available on Hugging Face, enabling fine-tuning for custom styles or characters.
- ComfyUI integration allows local execution, currently limited to 768p resolution.
- Commercially restrictive license targets smaller entities under $20M revenue.
- ByteDance launched the competing closed-source Seedance 2.5 on the same day.
Decoder
- Context-IR: An intermediate representation layer that parses prompts and reference media into a structured format for the generative model.
- ComfyUI: A node-based GUI for Stable Diffusion and other generative models, allowing users to build complex inference workflows.
Original Article
MiniMax releases H3 video model weights, putting an open model at the top of a video ranking for the first time. Artificial Analysis ranks H3 first in Video Editing, second in Text-to-Video, and third in Image-to-Video. The 33-billion-parameter model processes text, images, video, and audio together, generating four- to 15-second clips with stereo sound. According to the model card, a single prompt can include up to nine reference images, three video clips, and three audio clips.
Two pieces remain closed, though. The 2K resolution module and H3-Context-IR, which translates prompts and reference material into a structured intermediate format, aren't included. Running H3 locally in ComfyUI tops out at 768p, and users will need to handle context prep themselves using MiniMax's published prompting guides. The open weights do allow fine-tuning on custom footage, characters, or a specific visual style. One catch on the license side: commercial use is only permitted for companies making under $20 million in revenue.
ByteDance released its closed Seedance 2.5 the same day, which generates 30-second clips with built-in audio.
Google Workspace Plugins
Cursor's new marketplace plugins provide coding agents with native read/write access to Gmail, Google Drive, and Calendar.
Summary
Original Article
Changelog
Cursor can now read, write, and act across your Google Workspace.
New plugins give coding agents direct access to Gmail, Google Drive, and Calendar, so you can pull context, draft and update files, and manage your inbox and calendar without leaving Cursor.
Install plugins to connect:
- Google Drive: search files and folders, open and download content, create and organize files
- Gmail: search and read mail, draft and send messages, apply labels and manage threads
- Google Calendar: read schedules, create and update events, find free time
Browse the new plugins in the Cursor Marketplace or install them from the Customize page in Cursor. Learn more in our docs.
Mac Studio: Here's what Apple has in store for this fall and beyond
Apple is planning a major October update to the Mac Studio featuring the M5 Max and M5 Ultra chips to keep its workstation competitive.
Summary
Decoder
- Unified Memory: A high-speed memory architecture shared between the CPU, GPU, and neural engine, eliminating the need to copy data between separate pools.
Original Article
Apple is expected to refresh the Mac Studio this October with M5 Max and M5 Ultra chips, bringing significant performance improvements, up to 36 CPU cores, 80 GPU cores, and potentially 768GB of unified memory, though memory availability may be limited by supply constraints. The update also marks the arrival of the first M5 Ultra, as Apple never released an M4 Ultra chip, but higher prices are likely following recent Mac price increases. Looking further ahead, a 2028 Mac Studio with M7 Max and M7 Ultra chips is expected to deliver major AI performance gains, a redesigned cooling system, and support for up to 1.5TB of unified memory, making it Apple's most powerful Mac yet.
How to Make Your Design System AI-Ready
AI struggles with design systems because it cannot resolve the inconsistencies and undocumented hard-coded values that plague most modern projects.
Summary
Decoder
- Design Token: A single source of truth for design decisions (like color, spacing, or typography) stored as a machine-readable variable.
Original Article
AI-generated prototypes often fall short not due to model limitations, but because of undocumented decisions, hard-coded values, and inconsistencies scattered throughout design systems. A practical framework addresses this through treating design decisions as infrastructure, auditing tools like FigmaLint, and a three-layer system of spec files, token layers, and automated audits. Design debt can't be resolved by AI alone—it requires deliberate human guidance and ongoing maintenance of design systems over time.
AI Visual Production Platform (Website)
Caimera uses generative AI to automate fashion and e-commerce product photography, claiming a 99% cost reduction and 7.3x sales increase for users.
Summary
Deep Dive
- Automates flat-lay to on-model transformations and ghost-mannequin effects.
- Offers custom AI model training for brands to ensure consistent style.
- Provides API and bulk-processing workflows for enterprise DAM integration.
- Focuses on reducing production costs while maintaining high-resolution output for large-scale retail.
Decoder
- DAM (Digital Asset Management): Systems used to store, organize, and retrieve digital media like photos and videos.
- Flat-lay: A photography technique where items are arranged on a flat surface and shot from above.
- Ghost mannequin: An image editing technique that removes the mannequin from a piece of clothing, making it appear as if it is being worn by an invisible person.
Original Article
Full article content is not available for inline reading.
Professional Motion Studio (Website)
Premation is an open-source, WebGPU-based motion design studio built to provide deterministic exports and After Effects-like workflows.
Summary
Deep Dive
- Uses a deterministic rendering engine where frame index / fps ensures identical results across all hardware.
- Core engine leverages WebGPU for viewport and final render throughput.
- Supports industry-standard formats including Lottie, ProRes 4444, and VP9 + alpha.
- Includes advanced features like 2.5D cameras, skeletal rigging, and per-glyph text animation.
- Offloads heavy encoding to ffmpeg processes to ensure the UI remains responsive.
Decoder
- WebGPU: A modern web standard for high-performance graphics and computation, providing lower-level access to the GPU than WebGL.
- 2.5D: A technique where 2D objects are placed within a 3D space, allowing for perspective, depth-of-field, and rotation effects.
- Deterministic: A system that produces the same output for a given input every time, regardless of background noise or system timing.
Original Article
The WebGPU Motion Studio.
Compositions, bezier keyframes, 2.5D cameras, vector rigging, and deterministic exports in an open-source studio.
Supported formats: SVG, Lottie, MP4 - H.264, WebM - VP9 + alpha, MOV - ProRes 4444, PNG sequence, GIF, Image sequences, WAV audio
Engineered for Motion Designers
Everything you need in one WebGPU pipeline. Keyframing, 2.5D cameras, lights, vector tools, and skeletal rigging.
Cameras, lights and depth
Make any layer 3D to gain Z position and X/Y rotation. One- and two-node cameras with depth of field, lights with cone angle and shadows, extrusion and bevels for shapes and text. Grab a camera or a light in an orthographic view and move it, parent it to a null, and animate from there.
Keyframes on any numeric property
Linear, ease, ease-in, ease-out, hold and custom bezier. The graph editor edits value and velocity curves directly, and position keeps X and Y as independent tracks so they can be eased separately.
Per-glyph text animators
A real selector stack: animate position, scale, rotation, opacity, tracking or colour across characters, words or lines, with range and wiggly selectors. Text on a path included.
Bones and puppet pins on flat artwork
The bone tool builds a skeleton with forward kinematics, linear blend skinning and FABRIK IK. The puppet tool deforms a triangulated mesh from pins with an as-rigid-as-possible solver.
Paths that behave
Pen, pencil, pressure-sensitive brush and curvature tools. Boolean merges, trim paths, pucker/bloat, twist, zig-zag, repeaters, dashed and tapered strokes.
A deterministic emitter
Rate, lifetime, velocity, gravity, turbulence, size and colour over life. The same field always produces the same field, so a re-render is identical to the first one.
Unified Timeline Engine
Every panel edits the same document model with instant undo/redo history.
Multiple compositions per project, each with its own size, frame rate, duration and background. Nest one inside another as a pre-comp with time remapping, and instance it more than once. A placed composition is sealed - it renders at its own size, through its own camera - until you collapse transformations and let it meet the host’s.
Compose
Multiple compositions per project, each with its own size, frame rate, duration and background. Nest one inside another as a pre-comp with time remapping, and instance it more than once. A placed composition is sealed - it renders at its own size, through its own camera - until you collapse transformations and let it meet the host’s.
Animate
Set a keyframe on any numeric property, then shape it. F9 for easy ease, Shift-G for the graph editor, U to reveal everything animated on a layer. Interpolation is per-keyframe, not per-property.
Stage in 3D
Push layers apart in Z, add a camera and move through them. Orthographic top / front / left views and multi-view layouts, with move, rotate, scale and universal gizmos.
Render
Frame time is exactly index / fps, never wall-clock, so two exports of the same project produce identical frames. ffmpeg encodes in a child process, so a long export leaves the app usable.
Natural Language Keyframing
Describe the motion prompt to keyframe values or offset tracks. Every action applies as a single reversible edit.
Deterministic Frame Rendering
One WebGPU frame loop feeds every format. Frame index divided by FPS ensures identical exports every single time.
-
Deterministic by construction
One frame loop feeds every format, and frame time is index / fps rather than wall-clock. Re-render after a crash and you resume identical work.
-
It refuses to write an empty file
Every path asserts on a real frame count. A zero-frame encode is an error, not a silently successful black video.
-
Peak memory is one frame
Frames go to a temp directory and ffmpeg encodes in a child process. A multi-gigabyte export never passes through the app's memory.
-
Preview the actual export
The export dialog renders real export frames with the same snapshot builder the encoder receives, and lets you scrub the whole range first.
After Effects Muscle Memory
Maps directly to familiar shortcuts, tools, blend modes, track mattes, and pre-compositions.
- All 17 blend modes, track and alpha mattes, parenting and adjustment layers
- Classic-3D layer model: flat planes in 3D space, lit and shadowed
- Keyframe assistants - easy ease, time-reverse, sequence layers, stagger
- Expression controls: slider, angle, point, colour, checkbox, dropdown, layer
- Collapse Transformations and Continuous Rasterization, on the same sunburst column
- Interpret Footage: conform frame rate, pixel aspect, alpha, loop - stored on the asset
- Proxies, generated or attached, substituted in the viewport and never in output
Open-Source Core, Optional Hosting
Run the full desktop studio locally for free. Cloud plans add synced projects and remote rendering.
Self-hosted
The whole editor on your own machine, under the AGPL.
- Every tool — compositions, graph editor, 2.5D, rigging, particles
- No account, no network — projects are .motion bundles on disk
- Export to every format, offline, with no watermark
- The AI assistant with your own provider key
- AGPL-3.0 source — read it, build it, fork it
Cloud
Everything in self-hosted, plus the cloud — free while we're in beta.
- Cloud projects, autosave and version history
- End-to-end encrypted project sync across machines
- Hosted mp4 rendering — your machine stays free
- The plugin registry, and update checks
- The AI assistant with your own provider key
Frequently Asked Questions
Everything you need to know about Premation Studio, WebGPU rendering, licensing, and After Effects project compatibility.
Start Animating in Seconds
Free public beta. Local JSON project files on your disk with zero sign-up required.
Opening it the first time
macOS blocks apps from unidentified developers, so the first open needs the right-click menu. Once, per copy of the app.
-
Drag Premation to Applications
Open the .dmg and drag it across. Take Premation-macOS-arm64.dmg on Apple silicon and Premation-macOS-x64.dmg on Intel. The arm64 build will not run on an Intel Mac at all.
-
Right-click it, then choose Open
Not a double-click. Control-click or right-click the app in Applications. The right-click route adds an Open button that a double-click does not offer.
-
Confirm Open in the dialog
That is the whole workaround. macOS remembers the decision for that copy, so normal double-clicking works from then on.
System Requirement Notice
Premation renders every frame on the GPU via WebGPU or WebGL2 fallback. A hardware graphics processor is required.
GPU & System Specs
Graphics Processor: Dedicated GPU (NVIDIA/AMD) or Apple Silicon (M1/M2/M3/M4). The engine picks WebGPU when it can and falls back to WebGL2; there is no software rasteriser.
Integrated Graphics: Intel Iris Xe / AMD Radeon Graphics with current drivers. Fine for lighter compositions. Note that a WebGL2 fallback can report a maximum texture size of 4096, which bounds how far Continuous Rasterization will re-raster vector content.
System Memory: 8GB RAM minimum / 16GB+ recommended. Memory scales with composition resolution and cached pre-comps. Export is the exception: frames stream to disk one at a time, so peak export memory is one frame rather than the whole render.
Operating system: Windows 10 / 11 x64 & macOS 13+ (Apple Silicon & Intel). A desktop application running in Electron with native menus and filesystem access. Building from source needs Node.js 20 or newer; ffmpeg is optional for MP4, WebM, GIF and ProRes export.
SpaceX is set to acquire 130,000 acres of marshland in southern Louisiana
SpaceX is reportedly finalizing a deal to acquire 130,000 acres of Louisiana marshland for a new, strategically located launch site.
Summary
Decoder
- Polar orbit: An orbit that passes over the north and south poles, allowing satellites to scan the entire surface of the Earth as it rotates underneath.
Original Article
SpaceX and the state of Louisiana are close to finalizing a deal for the launch company to acquire about 130,000 acres along the northern coast of the Gulf of Mexico.
There have been persistent rumors about such an agreement for months, but now The Times-Picayune | The New Orleans Advocate reports that Louisiana Gov. Jeff Landry is expected to announce the agreement later this month.
The deal would give SpaceX control of an 18-mile stretch of marshland southwest of Lafayette. The site, known as Pecan Island, became available as part of a legal settlement that resolves dozens of lawsuits that blame ExxonMobil for pollution and coastal land loss, the newspaper reports.
No Louisiana officials publicly commented on the deal. Nor did SpaceX. In May, however, the company said in response to rumors about the Louisiana site on X, “It’s no secret that we intend to launch Starship a lot, targeting thousands of flights per year. That cadence will require the ability to launch from many different locations, so we are constantly exploring to find viable sites to expand Starship operations in the future, both domestically and internationally.”
Plenty of incentives
Adding fuel to the rumors was the passage of bills by the Louisiana Legislature earlier this year that included a package of incentives for aerospace companies, including liability protections and property tax breaks.
According to the Louisiana newspaper, the agreement with SpaceX will include provisions for coastal restoration and preservation of the sensitive marshland along the northern Gulf Coast, which is important for wildlife and also acts as a buffer for hurricanes that regularly impact the region.
So why would SpaceX be interested in a remote, marshy location in southern Louisiana?
Lots of reasons to like Louisiana
If the company is to fulfill its ambitions to launch thousands of Starship rockets a year to build a massive constellation of orbital data centers, among other purposes, it needs more launch sites. And there are limited expanses of undeveloped coastal locations along the Gulf of Mexico and southern Atlantic Ocean in the United States.
In the near future, SpaceX will have two orbital launch towers at its Starbase facility in South Texas and two more in Florida at Cape Canaveral. However, both of these locations are largely built out or congested with other launch companies.
The Louisiana location, close to the Intracoastal Waterway with deep-water access and far more available real estate than Starbase, would offer several significant advantages.
It is relatively close, by barge, to SpaceX’s massive Starfactory in South Texas near Boca Chica beach. Additionally, the southward-facing location offers potential access to polar orbits. In contrast to equatorial orbits that move from west to east, polar orbits go north to south (or vice versa) and generally pass near or over the poles. It is likely that a majority of SpaceX’s orbital data center satellites will wind up in near-polar orbits. From Louisiana, it is possible that a Starship could reach a polar orbit with only a short traverse over Mexico nearly 1,000 miles down range.
Hello, Henry Hub?
Another advantage of Louisiana is its proximity to natural gas infrastructure. For rapid launch operations, SpaceX will require extensive amounts of methane, and the Pecan Island site is located only a few dozen miles south of the “Henry Hub” distribution hub for natural gas, one of the most interconnected locations in the world. The availability of propellant would be far greater there than in Boca Chica or Florida.
A launch site in this area would raise significant environmental and logistical concerns and disrupt the local community. A similar thing has happened in South Texas over the last decade with the Starbase facility there, raising local opposition for environmental and other reasons. But from an economic standpoint, the Texas site has been good business for the state. SpaceX employs about 5,000 people directly in the Brownsville region and has transformed a sleepy border area into an aerospace powerhouse. Some analyses show the site now supports up to 24,000 indirect and direct jobs. It has also increased tourism during launches.
In terms of accommodating SpaceX, Louisiana would certainly be incentivized by the potential for similar economic activity.
Base Power raises $1B to roll out its giant new home battery
Base Power has raised $1 billion to launch a 39.2-kWh home battery system that is exclusively available via a company-managed subscription model.
Summary
Decoder
- IP67: An Ingress Protection rating indicating the device is dust-tight and can withstand immersion in water up to 1 meter deep for 30 minutes.
Original Article
Base Power's Core is a 39.2-kWh home battery that can keep a house running for up to 36 hours. Core, which can be installed in under an hour, can be connected to solar and has a built-in port for recharging from a portable generator. The lithium iron phosphate battery is designed to operate between -22F and 122F and has an IP67 rating for submersion in up to 3 feet of water. Core will only be available as part of a subscription service, and the company will retain ownership of the battery.
LLMs reward expertise
Domain expertise is the most critical factor in effective prompting, as deep knowledge allows users to steer AI models toward superior, contextually aware outputs.
Summary
Original Article
In the 2010s, if you had technical gaps (say, you couldn’t write CSS), you had to either rely on a skilled colleague or just hope that the answer to your exact problem was out there on the internet. Today, everyone can write sort-of-okay CSS by delegating the task to an LLM. LLMs make everybody into a generalist.
Because of this, lots of people don’t think there’s any skill involved in working with LLMs. If you want the product that LLMs can deliver — PhD-level mathematics, pretty good but sometimes tasteless computer code, or awkward LinkedIn-style writing — you can simply ask for it. Since everyone is talking to the same models, “skilled prompters” are getting the same results as people touching LLMs for the first time.
This is wrong. The most important skill in prompting is expertise in the domain you’re prompting for.
A good illustration of this is Terence Tao’s conversation with ChatGPT about the recently-discovered counterexample to the Jacobian Conjecture. This is not the same ChatGPT I talk to! I couldn’t get to where Tao gets, even with unlimited tokens to burn.
There’s a lot to learn about good prompting from Tao’s conversation. Here are a few observations:
- Tao’s messages are very short and to-the-point. He doesn’t respond point-by-point to the model, just to the gist
- The model outputs are much more concise than when I try and talk to GPT-5.6 Sol about mathematics. By signalling expertise, Tao shunts the model into “talking-to-mathematicians” mode, not “explaining-to-amateurs” mode
- Tao pushes back when the model’s responses look wrong, but he doesn’t directly contradict; instead, he says things like “this looks more complex than I was hoping for”
- Tao makes several leaps and suggestions himself. He almost never takes the model’s advice about where to go next
However, you can’t prompt like Tao on mathematical questions just by following these tips. The key to his technique is actually understanding the mathematics: pulling the relevant idea out of ChatGPT’s multi-paragraph response, suggesting alternate approaches or formulations, and identifying what “looks weird”.
Terence Tao is a better mathematician than I am a programmer. But the idea here — that domain knowledge makes you better at using LLMs — is something I’ve also experienced in my own work. If you have a good theory of your codebase, you can push the LLM much harder than if you have no familiarity. Because you have your own sense of what a good solution might look like, you can say “no, I think it could be simpler here”, or “but don’t we already do X?”, or “can we express this problem in these familiar terms?“.
This touches on an idea I’ve written about before: that system design problems are dominated by concrete specifics, not generic principles. Of course both are useful, but I’d rather have familiarity with the codebase than a deep general understanding of software systems. In his conversation, Terence Tao asks a lot of specific questions like “does X work here?”, or “given Y and Z, why A?“. I can’t ask those questions about the Jacobian Conjecture, but I can ask them about the systems I own at GitHub.
If you have no domain knowledge, you can cling onto the LLM to at least get something. That’s not bad! But if you have domain knowledge, you can wring far more value out of the same LLM by steering it hard in the direction you want. Most of us will have to do a mix of both these approaches, since we have domain knowledge in some areas but not others.
The usefulness of domain knowledge suggests that human expertise will continue to be useful even as models get stronger. For many tasks, the human is the bottleneck, not the model, because the difficult part is in communicating to the model exactly what kind of solution the human wants. The information is “in the model” already, but it takes a very smart human to pull it out.
What Are Companies Getting for All That AI Spending?
The Linux Foundation has launched the Tokenomics Foundation to standardize AI usage metrics as companies struggle to quantify the ROI of massive AI spending.
Summary
Decoder
- Tokenomics: The study or management of costs and efficiency associated with the production and consumption of AI tokens.
Original Article
The cost and benefits of AI are difficult to weigh. There are now thousands of models, and there is no consensus around how much energy a token requires, which model is best to perform a given task, or how many tokens it should take. The Linux Foundation established the Tokenomics Foundation in June to create common parameters for AI providers to disclose. Being able to measure the costs and benefits of AI will help researchers understand how AI is going to change everything.
Does Forecasting Have Room At The Top?
AI superforecasters are rapidly approaching human-level accuracy, potentially offering a 4–12 percentage point improvement in prediction market accuracy.
Summary
Deep Dive
- AI performance in forecasting is rising rapidly, approaching the skill level of top human superforecasters.
- Prediction markets are compared to simple statistical models to establish a baseline for 'miraculous' versus 'non-miraculous' gains.
- The 'Non-Miracle Rule' suggests we should expect incremental, not revolutionary, progress in predictive accuracy.
- Geopolitical forecasting involves significant irreducible uncertainty, unlike highly optimized domains like chess.
- Potential gains from AI in prediction markets are estimated at 4-12 percentage points.
- AI could excel at identifying 'unknown unknowns' that human analysts might miss due to cognitive biases.
Decoder
- Superforecaster: An individual who consistently makes accurate predictions about world events, often using structured probabilistic thinking.
- Brier Score: A scoring rule that measures the accuracy of probabilistic predictions, where a lower score is better.
- Aleatoric Uncertainty: The inherent, irreducible randomness in a system that cannot be explained or predicted, even with perfect information.
Original Article
Does Forecasting Have Room At The Top?
Superforecasting is the art/science/sport of predicting the future - for example, who will win elections, which countries will fight wars, when key technologies will be discovered. Over the past few years, it went from an obscure academic subfield to a multibillion dollar industry in the form of prediction markets. More recently, AI superforecasters have come close to the accuracy of top humans, and their performance is rising rapidly. In a year or two, we’ll see one of the following patterns:
Either humans have already come close to some fundamental limit on the predictability of world events - in which case AIs will plateau at or slightly above the human level - or the trend line will continue until AIs are far beyond top humans. By analogy to superintelligence, the natural term for the second situation would be “superforecasting”; since that’s already in use, we can cringely call it “ultraforecasting”.
Daniel Reeves makes the case for scenario A here. He describes a study he coauthored in 2010, which found that, on a variety of questions related to sports games and movie box office receipts, prediction markets only outperformed simple boring statistical models by 3-6%. Maybe those statistical models are close to the best that it’s possible to do; the rest is what the mathematicians call aleatoric uncertainty - irreducible complexity downstream of chaotic systems that entirely resist modeling.
He could be right. This post isn’t meant to be a decisive refutation, but rather a description of why I’m still about 70-30 expecting Scenario B.
Slightly Contra Goel, Reeves, et al
Reeves’ study claims that the prediction markets of 2010 only beat dumb statistical models by 3% (for sports) to 6% (for movies).
But these percentages aren’t real win-loss percentages; they’re variation in a quantity called root mean-squared error. One way to get a feel for this quantity is that the dumb statistical model for sports (home team advantage + win-loss record) beat an even dumber statistical model (home team advantage only) by 0.8 percentage points, and the prediction market beat the first (better) statistical model by another 0.4 pp. So the effect of going from a statistical model to a prediction market is half as large as the effect of knowing which two teams were playing and how good they are! On this metric, the prediction markets are a vast improvement over previous state-of-the-art.
Why can we frame this same result as either very small or very large?
Sports are optimized against prediction. If there were a fully predictable sport (eg heightball, where all athletes line up in a row and the tallest one wins), nobody would watch it. Instead, we go to absurd lengths to keep the outcome uncertain. Salary caps, draft systems, etc try to ensure that all teams have exactly equal talent. Commercial incentives and ceiling effects ensure that they have exactly equal training. Then an exactly-equal number of these exactly-equally-talented-and-trained people are placed in exactly-identical positions on a perfectly-symmetrical field and told to hit/kick/throw a ball which is placed exactly equidistant between both of them. It’s funny for me to describe it this way, because obviously this is what we want (to “keep things fair”), but it’s all designed for prediction-resistance. Given the difficulty of the domain, even very large relative advances in prediction look small in absolute terms.
Maybe we should look at Reeves’ other example, movie box office receipts. Here the markets did slightly better, getting a 6% improvement. But isn’t this still low?
Box office receipts differ by orders of magnitude (some movies make $100,000, others make $100 million), so the paper puts this on a log scale. 6% improvement on a log scale is already starting to sound pretty good.
And again, it all ends up coming down to what we compare it to. Here the super-dumb model is that all movies make $8.1 million, the takings of the exact average movie. The slightly-less-dumb model then adds the number of screens that the movie is showing on and the amount of Google search traffic for the movie! For example, a random indie film might be showing on three screens in the entire country, and a Disney blockbuster might be showing on ten thousand. The challenge the paper gives prediction markets is to significantly improve on knowing whether a film is an indie film or a Disney blockbuster, plus knowing how many people are interested in seeing it, and it has to do this on a log scale! No wonder the relative improvement number comes out looking slightly anemic.
I would summarize this section as: we shouldn’t expect miracles, but this doesn’t rule out further normal-sized gains.
Applying The Non-Miracle Rule To Geopolitics
Let’s return to the picture we looked at above. This graph is loosely based on Metaculus’ AIs vs. humans results: and these are mostly on geopolitical questions. What can Reeves’ model tell us about these?
In one sense, it can already be proven not to apply. In the sports analysis, the spread between base rate (guessing 50% on everything) and smart humans (the prediction markets) was 0.04 Brier score points.
But on Metaculus, the spread between base rate and smart humans (Metaculus Pro Forecasters) is 0.13 Brier points. So Metaculus’ superforecasters have already removed 3x more uncertainty than a naive port of Reeves’ model would suggest exists!
Here we return to the idea of sports as a uniquely unpredictable domain. To give a trivial example, the easiest sports question (will the best team in the league beat the worst team in the league?) might still only be 90-10 (upsets happen). But the easiest geopolitical question (maybe “will the US bomb Canada in the next month?”) could easily be 99-1 or more. So it’s unsurprising that geopolitics forecasters easily beat Reeves’ supposed upper bound.
What would a non-naive port of Reeves’ model say about geopolitics, where we try to import broad lessons rather than specific numbers? We might conclude that prediction markets somewhat improve on other state-of-the-art forecasting methods in determining whether the US will bomb Iran this year. But they would find that the magnitude of that improvement was relatively low compared to knowing other things, like whether the country we’re talking about is Iran or Canada.
Yes, this is dumb and trivial, which is a corollary of the original sports results being dumb and trivial. Again, all we’ve learned so far is not to expect miracles.
There might be more interesting variation here. Consider a question like “Will AI take most human jobs by 2050?” Here it seems like if you were a supergenius who really understood AI, you could predict whether scaling laws would continue to hold, and have a good sense of whether AGI is within easy reach. And if you were a supergenius who really understood economics, you could model whether cheap AI labor would be a complement or substitute to human work. So a question like this one might be the opposite of a sports game, and have almost no aleatoric uncertainty - maximally smart predictors could know almost everything that there is to know about it.
For the rest of this post, we’ll ignore all of this and think of geopolitical forecasting as a formless mass easily represented by summary statistics.
Anchors For Possible Non-Miraculous Gains
What would “non-miraculous” progress on Metaculus’ graph look like?
The vertical axis’ “forecasting score” is an Elo-like construct with GPT-4’s forecasting ability fixed as zero, and the following other reference points (all approximate, some guessed via napkin-math):
- Random chance (50% on every binary question): -20
- Average human: 0
- Best out-of-the-box AI: 20
- Wisdom of crowds: 25
- Best AI superforecaster: 30
- Professional human superforecasters: 35
- Various prediction markets: 25 - 50?
- World’s best forecasting team, under optimal conditions: 40 - 50
- Perfect score: 100
How can we predict the logical limit on forecasting? Although we can’t do this convincingly, in the spirit of Cotra 2020, let’s set some “anchors” and see where they land us.
Anchor 1: What if best human → AI superforecaster was the same level of advance as dumb model → prediction market?
Reeves’ implicit point is that gains from seemingly-exciting new technologies are small (in his study, 3-6% of RMSE). What if the advantage of an AI over a human forecaster was of the same magnitude?
We can convert Reeves’ RMSE into something called a Brier Score, and then convert the Brier Score into the Metaculus Score above. When we do that, we find that AIs would max out at a Metaculus score of 56.
Anchor 2: What if best human → AI superforecaster was the same level of advance as professional superforecaster → Samotsvety?
In 2010, we probably would have guessed that 35 (“various professional superforecasters”) was near max possible forecasting performance.
Since then, more people have joined the forecasting world, allowing us to find superstars who perform better than the previous stars, and the community has produced some tools and algorithms to support their work. Samotsvety, a team of well-practiced superstars using good tools, consistently beats “ordinary” superforecasters. If AIs beat Samotsvety by the same amount Samotsvety beat the 2010-era SOTA, according to Samotsvety’s own (generous) accounting, then they would max out at a Metaculus score of 65.
Anchor 3: What if best human → AI superforecaster was the same level of advance as human → AI chessmaster?
AI first beat the human champion at chess in 1997. Since then, it’s continued improving, and is now far beyond any possibility of human catch-up.
This alone doesn’t help us much. We want to know how close to optimal chess play top humans are. If they’re 99% of the way, but AIs are 99.9% of the way, that still could look like AIs beating humans every time.
But we can translate this into an objective space by looking at handicap. When AI first beat humans, its lead was fragile; a one-pawn handicap was enough to restore human supremacy. Since then, it’s been gradually getting better, and is now up to a 2.4 pawn handicap. It’s been a bit slower in achieving the third pawn than in the first or second, but not too much slower, and it’s too soon to say with certainty that it’s starting to plateau.
So it seems top human play is at least 2.4 pawns (let’s round this off to 3 pawns, since we have no reason to think we’ve reached the plateau) short of optimum. What would it mean if top human forecasters were “three pawns” short of logical limits?
A three pawn handicap in chess is about 3x the difference between the world champion and a average grandmaster. If we round that off to 3x the difference between Samotsvety and Metaculus Pro Forecasters, the AI tops out around 85.
So What?
We sure have made a lot of claims about Metaculus Peer Forecasting scores, a number that approximately nobody cares about. How much real-life performance gain should we expect from max-possible AIs?
Consider a prediction market, currently at 50%, dominated by players somewhere between Metaculus Pro Forecaster-level skill and Samotsvety-level skill. And suppose that the market was about an event which would, in fact, happen. How confident would each of these max-possible-AIs, on average, be on the market?
Score 56: 54%
Score 65: 56%
Score 85: 62%
So these max-possible AIs would improve the accuracy of prediction markets by 4 - 12 percentage points.
This might sound disappointing - moving from 50% chance to 54% chance is hardly Nostradamus-level prescience. But I would actually be very excited by this. When I check something I care about on Polymarket, I usually have a decent sense of how likely it is. I’m not sure whether there’s a 55% chance or a 75% chance of Anthropic being worth more than OpenAI next year, but I know it’s not 10% or 90%. I visit Polymarket to see whether it’s more like 55% or 75%. On that metric, improving by 4% - let alone 12% - is actually pretty exciting.
I find these numbers more interesting than the worn-out debate over whether AIs or humans have reached the “max” and whether “perfect forecasting is impossible in principle”. It should be uncontroversial that the max is somewhere above current levels, even if only a fraction of a percent (for example, because even the best human forecaster sometimes has a bad day). And it should be uncontroversial that even the best possible forecaster can’t get 100% on everything (because you could ask questions about quantum effects that are unpredictable even in principle). Now that we’ve demonstrated that the answer will be some particular number, we can get to debating what number it is. If AI gives gains similar to past technologies, it could be between 4-12 %pp on prediction markets. Of course, if AI is much worse or better than past technologies, it could be some totally different amount.
Most people don’t have an intuitive sense of what 4-12 %pp on prediction markets buys you. I admit I sometimes blend together how confident I should be in an estimate by a professional superforecaster and how confident I should be in an estimate by Samotsvety, even though that’s a similarly-large difference. Probably all of this pales into insignificance compared to the gains we could get by switching from our current strategy of making decisions based on vibes and ballroom-related-bribery to listening to markets and forecasters at all.
And that 4-12 %pp is only the headline number. A forecaster who is that much better at prediction markets would also be better at crafting good policy, or at identifying unknown unknowns (eg spotting the risk of a coronvirus-style pandemic even if this was otherwise so far off the radar that nobody had thought to write a prediction market question about it). Smarter important-question-identifiers and smarter forecasters together would provide a joint advantage greater than the benefit of either alone, and smarter policy-crafters could compound on this by turning that knowledge into useful action.
All of this keeps me excited about AI superforecasters, even though I don’t expect miracles.
Ori Eval (Website)
OpenRouter has released Ori Eval to help developers benchmark and select the best models for their specific use cases.
Summary
Original Article
OpenRouter's Ori Eval helps users find the best model for what they're building.
Apple Plans iPhone-to-Windows Copy and Paste in EU After Microsoft Request
Apple is developing iPhone-to-Windows cross-device copy-paste functionality to comply with EU Digital Markets Act requirements.
Summary
Deep Dive
- The feature will allow seamless copy-paste between iPhone and Windows in the EU.
- It uses AccessorySetupKit, requiring explicit user consent for clipboard access.
- This is a significant deviation from Apple's typical cross-device integration strategy, which is usually exclusive to the macOS ecosystem.
- The implementation process is complex and will involve a long developer beta period.
- While currently limited to the EU, Apple has previously expanded similar DMA-compliant features globally.
Decoder
- Digital Markets Act (DMA): An EU regulation aimed at ensuring fair competition by requiring large 'gatekeeper' technology platforms to open their systems to interoperability with third-party products.
- Universal Clipboard: An Apple feature that allows users to copy text, images, or files on one Apple device and paste them onto another via iCloud.
Original Article
Apple Plans iPhone-to-Windows Copy and Paste in EU After Microsoft Request
Apple is working on a new feature that will support cross-device copy and paste between iOS devices and Windows PCs.
Microsoft asked for the feature using Apple's EU interoperability request system for developers, which Apple implemented to comply with the Digital Markets Act. Apple began evaluating the request in March, and on June 26, proposed a project plan.
Microsoft described how the feature would be used:
- Copy text or other supported content on their iPhone and paste it directly on their Windows PC, and vice versa.
- Perform common productivity workflows without needing to foreground an app or take explicit repeated actions to initiate a transfer.
- Experience clipboard synchronization as a lightweight, continuous capability rather than a manual, app-driven operation.
Apple proposed a solution that would let an iPhone share and import items from the pasteboard to a paired accessory, like a Windows PC. Apple plans to use a solution similar to the Accessory Notifications option it added for third-party wearables in the EU in iOS 26.5. Users would need to give a paired device one-time permission to paste content from an iPhone.
To adopt this solution, you will need to adopt AccessorySetupKit in order for the user to consent to the necessary permissions allowing the pasteboard to be shared with your accessory. This solution will require the user to perform a one-time consent to sharing the pasteboard with each paired accessory that will receive this functionality.
Right now, an app can only read an iPhone's clipboard if the app is actively running in the foreground and the user grants clipboard access with each copy/paste action. For its own devices, Apple has a Universal Clipboard feature with cross-device copy and paste between iOS and macOS.
Microsoft complained that the clipboard restrictions imposed on apps prevent the continuous synchronization of clipboard content from an iPhone to a PC or other connected device in the background.
Apple says introducing cross-device copy and paste is a significant engineering effort, with work expected to be complete by fall 2027. It will be shipped first in a developer beta for testing ahead of a public release. If the feature is added toward the end of fall 2027, it could be early 2028 before the public gets access.
It is likely that cross-device copy and paste will be limited to iPhone and Windows users in the European Union, though it is possible Apple could implement it worldwide.
Apple has brought some features it was required to support under the DMA to all countries, like eSIM transfers from iPhone to Android and simpler data transfers when switching between Android devices and iPhones. Other features, like AirPods-style pairing and notification forwarding for third-party wearables, have only rolled out in the EU.
Why Silicon Valley is divided over China's powerful, cheap AI models
The US tech industry is bitterly divided over whether Chinese open-weight AI models are essential low-cost tools or dangerous national security liabilities.
Summary
Deep Dive
- Moonshot's Kimi K3 model ranks high on intelligence indices, challenging US-based frontier models.
- Nvidia, Microsoft, Google, and Meta have publicly backed the importance of open-weight models in an open letter.
- Critics argue that open-weight models are prone to 'distillation attacks,' where Chinese firms train models using data scraped from US proprietary AI.
- Security concerns include potential backdoors and the inability to enforce safety guardrails on downloaded models.
- Some experts suggest adopting a 'movie classification' model for AI safety rather than outright bans.
- The Chinese government has accused the US of 'AI hegemonism,' framing their own open releases as a global public good.
Decoder
- Open-weight model: An AI model where the learned parameters (weights) are released publicly, allowing developers to run the model on their own infrastructure rather than relying on a closed API.
- Distillation: A process where a smaller, more efficient model is trained to mimic the outputs and behaviors of a larger, more powerful model.
- Distillation attack: The controversial practice of using a high-performing proprietary AI to train a competing open model, effectively 'stealing' the intellectual property and capability of the original.
Original Article
The surge of powerful Chinese open-weight AI models has sparked a feud in Silicon Valley and Washington over whether they are threats to U.S. national security or essential parts of the expanding AI economy.
Entrepreneurs and politicians are taking sides — arguing for either unfettered access or a strict crackdown — by forming alliances, issuing open letters, and launching social media volleys.
Here’s what you need to know.
What’s driving the Chinese open model debate?
While the leading American AI companies have kept their best models closed-source, Chinese AI labs have been releasing open-weight models as a way to attract global customers. These open models are becoming increasingly powerful.
Moonshot’s open-weight model, Kimi K3, now ranks fourth on Artificial Analysis’ intelligence index, behind only Anthropic’s Opus 5 and Fable 5 and OpenAI’s GPT-5.6 Sol.
While the adoption of Chinese models is on the rise in the U.S., politicians and companies have also sounded alarms. Some accuse Chinese companies of building powerful models with smuggled chips and distillation — using the output of American models to train their own. Others warn of the security risks posed by the powerful, free-to-download models.
Whether to embrace or reject these Chinese models has become a divisive issue in the U.S. tech industry.
What makes Chinese open models so competitive?
Chinese open models are cost-effective. Open-weight models are free to download, but running them incurs computing costs. Many users pay to access the models through their developers or third-party cloud providers.
Facing ballooning AI costs, some U.S. companies and developers have been diverting simpler tasks to Chinese models to save money.
Companies can also use Chinese open-weight models to build customized AI systems, or host the models locally so they don’t have to share data with external parties.
Why is Silicon Valley split?
Chinese open models benefit AI infrastructure providers and companies that need affordable AI services.
Nvidia has been a leading voice in support of open-weight models. The availability of open models creates a bigger customer base for its AI chips. On July 24, Nvidia chief executive Jensen Huang posted on X for the first time to share an open letter stressing the importance of open models in expanding access to the AI economy. Other tech giants, including Microsoft, Google, and Meta, have since joined the letter.
OpenAI, which has released open-weight models that are smaller than China’s models, also signed on. “i want the US to win in AI both in open source and proprietary models,” chief executive Sam Altman wrote in reply to Huang’s post on X.
i want the US to win in AI both in open source and proprietary models, and i am glad to see this
A group of 179 Silicon Valley startups also sent a letter to the Trump administration, calling on officials to preserve their access to open models.
On the other side, Anthropic has called for more restrictions on Chinese AI to protect America’s lead. In a statement, Dario Amodei argued that China could use its models to achieve military superiority or to repress its people. He added that open-weight models could also be misused to carry out cyber or biological attacks, because guardrails are hard to apply to them.
Flo Crivello, the founder of AI assistant startup Lindy, wrote on X that although his company used DeepSeek, he supported a ban on Chinese models to make sure America leads the global AI race and to prevent the “artificially cheap models” from hurting the U.S. AI ecosystem.
What are U.S. government officials saying?
Tech policymakers in the Trump administration have expressed different views on whether the government should crack down on Chinese AI.
Among the hardliners, Michael Kratsios, a top technology adviser to President Donald Trump, accused Moonshot of distilling Anthropic’s model Fable 5 and acquiring banned Nvidia chips. Treasury Secretary Scott Bessent has warned the U.S. could sanction Chinese companies for conducting “distillation attacks.”
On the other side, David Sacks, another adviser to Trump, has argued that restricting Chinese open models would undermine the competitiveness of the U.S. companies that use them. Commerce Secretary Howard Lutnick is also concerned about companies’ access to cheaper AI, according to a Politico report.
What about security risks?
U.S. officials have raised concerns about Chinese censorship and potential “backdoors” — hidden mechanisms that could trigger malicious behavior when a model encounters a specific input. Chinese models commonly censor content Beijing deems sensitive, though no security backdoors have been publicly documented in major Chinese models.
Ryan Fedasiuk, a fellow at the think tank American Enterprise Institute, said it is possible to address these security issues without banning the models. Potential solutions include post-training the models to remove censorship behaviors and labeling AI products built with Chinese models, Fedasiuk told Rest of World.
Both open and closed models can lead to cyberattacks. In July, an AI agent powered by OpenAI models acted on its own and broke into the AI firm Hugging Face’s infrastructure. Hugging Face turned to Chinese open model GLM-5.2 to analyze the attack, because the attack commands had been blocked by frontier AI models’ safety guardrails.
Hussein Abbass, an AI and computing professor at the University of New South Wales, Canberra, said that instead of a blanket ban on open models, the global community should develop systems to evaluate their safety risks and regulate model releases. “What comes to my mind is movie classification,” Abbass told Rest of World. “Every time a movie gets released, every country has its own classification for appropriate use.”
Proponents of open-source say open models can help fend off attacks.
On July 27, Nvidia announced a coalition of tech companies called the Open Secure AI Alliance, which promises to share open cybersecurity defense tools. The announcement called on regulators to see open models “as defensive assets, not liabilities.”
How is Beijing playing the narrative?
The Chinese government has positioned the country as a champion of an open, inclusive global AI order. Speaking at an AI conference last week, Chinese President Xi Jinping said China’s AI industry would provide a public good for the world, and implicitly criticized the U.S. for placing its own national security above that of others.
On July 27, China’s Ministry of Commerce also accused the U.S. of “AI hegemonism” over its distillation allegations, saying American companies had also distilled Chinese models. The ministry threatened countermeasures if Chinese interests were harmed.
Anthropic pays AI's biggest salaries. Its CEO just discovered people might take them for the money
Anthropic is reportedly the highest-paying AI lab, yet CEO Dario Amodei finds that high salaries fail to guarantee researcher loyalty or mission alignment.
Summary
Original Article
The most valuable workers in technology cannot be kept, whatever the price. That is the uncomfortable admission running through the AI industry this week, as elite researchers keep hopping between labs that have offered them almost everything money can buy.
The trigger was a single line in an Axios report on the talent wars. A source said Anthropic’s chief executive, Dario Amodei, has grown worried that new hires are joining for the money rather than the mission. The internet did the rest.
Within hours the line was everywhere, mostly as a joke. “Breaking: people work for money,” ran one widely shared reply. “Local man discovers economy,” ran another. The engineer Gergely Orosz made the sharper point. Anthropic reportedly pays more than any lab in AI, even OpenAI. So fretting that people come for the pay is a strange complaint from the firm writing the biggest cheques.
The churn is real
Under the mockery is a genuine problem. Hiring a top AI researcher is hard. Keeping one is turning out to be harder. The news peg was Lilian Weng, a co-founder of Mira Murati’s Thinking Machines Lab, who walked out last week and resurfaced days later at OpenAI. She was the fourth Thinking Machines co-founder to leave in a year.
She is not alone. Google lost Noam Shazeer to OpenAI and the Nobel winner John Jumper to Anthropic in June. Meta paid heavily to stock Alexandr Wang’s superintelligence team, then watched prize recruits leave again, several for OpenAI. Frontier AI now behaves like one connected ecosystem, with the same few hundred people circulating through it.
The money is staggering, and it is not the whole story. Researchers also chase compute, influence over what gets built, and the freedom to work their own way. Many believe they are reshaping the world economy, which makes status and ideology weigh as much as the fortunes on offer. Some also sense a closing window, and want equity locked in before an IPO.
Why “mission” is the last lever
That is what lifts Amodei’s worry above a punchline. When every lab can pay millions, pay stops telling people apart. Mission is the only lever left, and it is the one money cannot verify. As one developer put it, nobody can really test conviction while the cheques are this large. What happens to all that principle if a valuation slips?
The strain shows in how labs now hire. Reports describe Anthropic culture interviews that probe whether candidates fear the technology enough, a screen critics mock as cult-like. More than 1,300 lab employees recently signed a letter warning that development could outrun control. The people building the systems are the ones most publicly anxious about them.
There is a quieter cost, too. Research runs on trust and continuity, and constant departures interrupt projects and force labs to pay twice for the same small circle. More than 400 former Apple staff now work at OpenAI, which Apple is suing over trade secrets. Poaching at this speed leaves lawsuits and half-finished work behind it.
The bottom line is the one Axios landed on, and the one the memes accidentally proved. These companies can buy a researcher’s time. Lasting loyalty is not for sale, and the labs paying the most are the first to learn it. That belongs, fittingly, to the mission they keep insisting the money is not about.
Former OpenAI exec Fidji Simo discusses her battle with POTS and her startup's plans to cure it with AI and 3,500 vials of blood
Former OpenAI executive Fidji Simo launched ChronicleBio to identify biological sub-diseases for chronic conditions using AI-driven analysis of 3,500 vials of blood.
Summary
Decoder
- POTS (Postural Orthostatic Tachycardia Syndrome): A condition characterized by an abnormal increase in heart rate upon standing, often resulting in severe fatigue, dizziness, and cognitive impairment.
- Mobile Phlebotomy: A medical service where specialized staff perform blood draws at a patient's home rather than in a hospital or clinic.
- Longitudinal Data: Information collected from the same subjects over an extended period to track changes or disease progression.
Original Article
Exclusive: Former OpenAI exec Fidji Simo discusses her battle with POTS and her startup’s plans to cure it with AI and 3,500 vials of blood (so far)
When Fidji Simo announced she was leaving her role as one of the most senior members of OpenAI's leadership team after a seven-year battle with her chronic illness, Postural Orthostatic Tachycardia Syndrome, or POTS, the news felt unexpectedly personal. Simo's post said she would be focusing on how to use AI to cure these types of diseases.
"Do people like my cousin have any reason to have hope that AI can actually make a difference for their health?" I asked Simo when I reached out to her after the announcement. Simo was previously at Meta for a decade, where she oversaw the Facebook app, and then served as CEO of Instacart, which she brought public in 2023, before joining OpenAI in 2025.
"Yes," she answered. "I created a company, ChronicleBio, to tackle just that."
We hopped on the phone to chat about it in Simo's first interview since leaving her position as OpenAI's CEO of AGI deployment, where she reported directly to CEO Sam Altman. While now her main focus is her recovery and a never-ending schedule of medical appointments, she's also working on growing ChronicleBio as well as continuing to advise OpenAI.
ChronicleBio's three cofounders—Simo, Rohit Gupta, and Rishi Reddy—also either have chronic diseases themselves, or have a family member with one. These days, Simo says she's "physically the worst I've ever been." There is no cure for POTS. It causes dizziness upon standing up, fatigue, brain fog, headaches, and other symptoms, owing to an imbalance in the body's autonomic nervous system.
Chronic conditions are "becoming a real epidemic," Simo tells me. "We're talking about hundreds of billions in lost productivity, and so there's very big potential in finding drugs for these conditions."
In its first year as a company, ChronicleBio has performed 890 blood draws from 709 patients in Utah, Arizona, Texas, and India. It has over 3,500 tubes of blood in its "biobank," the company tells me. It's extracted 153 terabytes of data from the blood—that's three times the 45 terabytes GPT-3.5, a 2022 model from OpenAI, was trained on. The company has raised $15 million to date.
The next big thing: home blood draws. On Aug. 11, ChronicleBio will launch a sign-up link for mobile phlebotomy trucks to come to the homes of people with certain chronic diseases. Participants will get an in-depth report on their condition, free for the first 250 people. In exchange, they'll give their biological data to ChronicleBio.
The goal is to learn more about diseases and improve the success of clinical drug trials, something Simo says would be nearly impossible without AI.
In preparation for this interview, you sent me an article that you said encapsulates ChronicleBio's approach. It talks about how some patients with long COVID were participating in a clinical trial. The drug was working well for them, but then the trial was canceled for supposedly being ineffective for the group as a whole.
Fidji Simo: Yes, so that's really what ChronicleBio is meant to solve. We have seen a lot of clinical trials fail because the pharmaceutical companies aren't able to identify which subset of patients [a drug] could work for. So they end up giving the drug to everyone with the same diagnosis. Let's say it's POTS. But there could actually be five sub-diseases within POTS, and the drug would work for one of them, but not the other four. So the clinical trial fails when it could have succeeded if we could have identified these people upfront. It seems really simple, but it hasn't been done for these conditions.
So what your company is doing is finding patients with similar symptoms, grouping them together, and then testing drugs on those subgroups so the trial is more likely to be successful?
That's exactly right. We have already found five sub-diseases where the biology is really different, despite the symptoms being the same. And now that we understand the biology, we can map that to existing drugs that would solve the problem, and so we're going to start testing these existing drugs on our patient population before the end of the year. Then we would partner with biotech and pharma to develop new drugs, with the goal of having suitable therapeutics for every part of this patient population.
What exactly do you mean by a sub-disease?
The sub-diseases don't even have names right now. That's the problem. So, the way the medical system names these syndromes is by their symptoms. In the case of POTS, it's called Postural Orthostatic Tachycardia Syndrome. It's basically named after the symptom: Tachycardia means your heart rate goes up when you stand. But for one group of patients the disease might be driven by the immune system. For another, it's driven by the mitochondria. The underlying biology is very different, and that's why one drug isn't going to work across everyone even if the symptoms are the same.
Very cool. Backing up for a second, is it an amazing feeling to have gone through such a long medical journey yourself, and now you're in a position of power to actually improve the system?
Yeah, you know, it's obviously a horrible disease, and I certainly wish I could have dodged it. But at the same time, I think it has given me enormous meaning. The delta between the disability from these diseases and the amount of funding and research being done on them is terrible. If you look at a condition like chronic fatigue syndrome, it is considered the most disabling disease of all diseases.
And yet, if you look at the amount of funding for this condition, it's absolutely pathetic for two reasons: One, it primarily affects women, so of course you get less funding. Second, while it completely disables you, it usually doesn't kill you. And so the combination of these two things has made it that these diseases are really ignored, even though they affect people at the prime of their lives.
I have a lot of empathy for people with chronic fatigue and chronic conditions after being pregnant. It kind of feels like that.
Yeah, imagine that 24/7, impossible to move. [Some] patients are fully bedridden in the dark, sensitive to light, sensitive to sound. It's a really terrible quality of life, and to me, it seems impossible that with the tools we have today, we would continue to conclude that diseases are incurable and that patients should be in a dark room for years. We owe them something better, given the progress that we're seeing in a lot of disciplines.
So what's different now with AI? What does it unlock that wouldn't have been possible before?
The complexity of these diseases made it that without AI they were incredibly difficult to solve. Like I said, they're multisystem, so you need to be looking at the state of the nervous system, the state of the immune system, and how it correlates with your genetics. All of that is a massive data problem that was very hard to get your hands around without AI. And so finally we have AI, and then on top of that, you have the cost of these analyses going down. Doing a genetic analysis years ago was way more costly than it is now. Analyzing 150 terabytes of data would have been either impossible or would have taken years, and now it takes us minutes.
So that's what gives me a lot of hope. I'm physically the worst I've ever been, but at the same time, we are at a moment in time where we have the best tools we've ever had to solve diseases that are considered incurable.
What AI models are you using?
We're using a combo of OpenAI and Anthropic models. We're using anything that's available that can help.
Why do you need to collect blood to get the right data?
The reason I did ChronicleBio is because I really think that we are missing true biological data to make progress towards discovering drugs. Right now, a lot of the models use a lot of EHR data—medical records. But medical records don't tell you enough about biology. They're incredibly noisy. They don't tell you how the human body works. If you look at LLMs, they work so well because the internet existed, right? You already had all of this language. We are missing the internet of biology.
What's the latest initiative you're working on?
Right now we've acquired all of this data [from blood] by partnering with clinics, but we think it's really important to get that data from anyone who wants to participate. We're now in the process of opening up our tests to anyone in the U.S., with mobile phlebotomy coming to their house. That's going to allow us to have a much larger dataset, but also reach patients that are bedridden, that are in the sickest stages of the disease. And we actually return the data to patients, so that gives them more information about that condition in case that can help direct them towards a particular therapy.
That's amazing. When does it start?
It's next week [on Aug. 11]. We partnered with mobile phlebotomy companies that collect the blood in a kit. They send that to us. We get it analyzed. It takes a couple weeks because these analyses are very robust. And then we send back a report to the patient about everything we learned, and that data goes into our database. And then over time, if we have more findings about which sub-disease the patient might have or things like that, we continue keeping them posted, and then they can take the test over time, so at multiple points in time, so that we can also see how they evolve.
There are already a variety of mail-in blood tests out there. How is what you're doing different?
That's right. The test we do is very focused on these particular complex chronic conditions. So it's not just the standard blood tests that are common. It's a really advanced research-grade blood test.
How much will it cost?
We're making it free for the first 250 patients because we really want to make sure they are getting value out of the report. After that, it's going to cost $400. We're doing it at cost, meaning that's what it costs us, and we're charging the same for patients. The whole point for us is not to make money. It's to collect data so we can find cures.
This is all so fascinating. I'm glad we did this.
Thank you for your interest! We're excited. You know, when I was at OpenAI, I said, "I think if AI accomplishes everything but doesn't cure disease, that would be a very sad state of affairs." The real promise of AI has always been to cure disease. I think it would be a tragedy if we had all of these amazing tools in our hands, but weren't able to turn them into drugs that can save patients' lives on a time frame that matters.
MirrorCode
MirrorCode tests if AI can rewrite complex, long-horizon software from scratch without original source code access.
Summary
Original Article
Scale-aware evaluations
Crucially, we provide a large enough inference budget to make a serious attempt at MirrorCode tasks. Many existing software engineering benchmarks limit inference spending to around $1–10, even when the task would take weeks for a human to complete. For example, one of the largest MirrorCode tasks cost $2,600 for a single run and involved AI working for 19 days without human intervention.
GPT-5.6 Sol Uses Twice the Tokens of GPT-5.5
OpenAI's GPT-5.6 Sol update effectively doubles the cost of Codex workflows by increasing token usage and introducing new cache-write fees.
Summary
Original Article
Full article content is not available for inline reading.
White House to host AI companies Tuesday to review new model-testing framework
The White House is hosting AI companies to review a voluntary framework that grants the government 30-day early access to frontier models.
Summary
Original Article
- The White House will meet with leading artificial intelligence companies Tuesday to review a completed voluntary framework for testing the cybersecurity capabilities of advanced AI models.
- Anthropic is expected to attend, while OpenAI and Google are also expected to participate, according to The Information.
- The framework would allow companies to give the government early access to certain frontier models for up to 30 days, but it cannot be used to create a mandatory licensing or preclearance system.
The White House will host artificial intelligence companies Tuesday to discuss a newly completed framework for reviewing the cybersecurity capabilities of the industry's most advanced models, a White House official confirmed to CNBC.
The meeting will focus on the voluntary framework President Donald Trump ordered in June, the official said, speaking on condition of anonymity to talk about the unannounced meeting. The Information first reported the planned meeting.
Representatives from Anthropic are expected to participate, according to a source familiar with the plans, who spoke on condition of anonymity to talk about the meeting. OpenAI and Google are also expected to attend, according to The Information. The White House official said the administration has been working with a broader group of industry partners.
Trump's June 2 executive order directed federal officials to create a process through which AI developers could determine whether models under development qualify as "covered frontier models."
Under the voluntary program, participating developers could provide the government access to those models for as long as 30 days before making them available to other trusted partners.
The administration has said the early access could help the government and technology companies evaluate whether powerful models could be used to discover software vulnerabilities or carry out sophisticated cyberattacks.
The order directed the Treasury Department, National Security Agency and Cybersecurity and Infrastructure Security Agency to establish a classified benchmarking process for assessing models' advanced cyber capabilities.
The benchmark and the threshold used to determine which models qualify for review are expected to remain classified. The White House has not publicly released the completed framework or detailed the metrics the government will use to test participating models.
The order explicitly states the program cannot be used to establish a mandatory federal licensing, permitting or preclearance requirement for the development or release of new AI models.
The framework arrives as leading AI developers increasingly test whether their systems can autonomously identify and exploit cybersecurity vulnerabilities.
Last month, OpenAI disclosed that an experimental AI agent escaped a restricted testing environment and compromised Hugging Face's systems while attempting to obtain answers for a cybersecurity evaluation.
Hugging Face CEO Clément Delangue told CNBC Monday the incident underscored the growing risks posed by increasingly autonomous AI systems.
DeepSeek's new AI model is by far the cheapest of well-known models to run, research firm says
DeepSeek’s V4-Flash model is significantly more cost-efficient than competitors, priced at 105 times cheaper than Anthropic’s Claude Fable 5.
Summary
Original Article
DeepSeek's V4-Flash AI model is the cheapest to run among well-known models, costing 105 times less than Anthropic's Claude Fable 5.
Google is working on Plugins for Gemini Enterprise
Google is developing 'Plugins' for Gemini Enterprise, likely enabling users to chain packaged workflows across company data services.
Summary
Decoder
- Connector: A service-level integration (e.g., Google Workspace, Microsoft 365) that enables an agent to authenticate and access external data.
- Skill: A reusable action or function that an AI agent can execute to perform a specific task within an enterprise workflow.
Original Article
Google appears to be preparing Plugins and a dedicated Notifications area for the Gemini Enterprise app, extending its push beyond chat toward reusable workplace workflows. Clues in an unfinished interface show a revised Connectors area split into three tabs: Connectors, Skills, and Plugins.
The Plugins tab is currently empty, indicating that the feature is still under development. Its placement suggests that Plugins may act as mini-apps or packaged workflows, combining reusable Skills with one or more Connectors. This could let business teams run multi-step processes across company data and services without configuring every component from scratch. The structure also resembles how Claude separates connectors and skills, while allowing packaged tools to build on both.
A separate Notifications area is also being prepared. It is expected to collect completed Gemini responses from scheduled tasks and conversations placed in a queue, giving users one place to return to work that finished in the background. This would be particularly useful for employees running research, reporting, or other jobs that take longer than a standard chat response.
The work fits Google’s broader strategy for Gemini Enterprise, which brings agents, company data, and workflow controls into one governed service. Google already describes Skills as reusable actions, Connectors as links to services including Google Workspace and Microsoft 365, and Inbox as a central hub for long-running agent activity. Plugins would add a packaged application layer on top of those pieces.
Google has not publicly announced the Plugins section or provided a rollout date. The empty catalog and unfinished Notifications interface indicate that both are still in development, with availability, supported accounts, and the first plugin partners remaining unknown.
Bringing Human Taste to AI Models
Design Arena, a feedback platform for AI models, has raised $7.9 million to quantify 'subjective taste' at a $60 million ARR run rate.
Summary
Deep Dive
- Platform generates $60M ARR by selling proprietary human-preference datasets to frontier AI labs.
- Business model revolves around A/B testing where users rank AI-generated outputs.
- Provides a vital training signal for models where 'quality' is subjective (e.g., game design, aesthetics).
- Distinguishes itself from failing competitors like Yupp by focusing heavily on enterprise-grade feedback loops.
- Data collection allows for geographic and temporal tracking of shifting aesthetic preferences.
Decoder
- ARR (Annual Recurring Revenue): A subscription-based business metric representing the total amount of predictable revenue generated in a year.
Original Article
As co-founder Grace Li tells it, her company started a few weeks before graduation in 2025, with a handful of college friends trying to make their AI game engine work. The models could make functional games, but none of the games were fun — which raised the interesting question, how can you tell if a game will be fun?
There was no substitute for human judgment, they decided, and soon they were brainstorming ways to get honest human feedback at scale. The result became Design Arena, an AI tool now used by 5.3 million people around the world. As it turned out, there were lots of AI companies looking for scalable user feedback — and many of them were willing to pay for it.
“It was the missing bottleneck for a lot of these models to make improvements in the design space,” Li says. “About a week later, we closed our first major deal with a frontier lab, and the rest is kind of history.”
On Monday, the company behind Design Arena — dubbed Intelligence — announced a $7.9 million seed round led by Index Ventures with participation from Conviction (Sarah Guo and Mike Vernal), A*, Valkyrie, and others.
For non-enterprise users, using Design Arena is a lot like using a sophisticated model router. There’s a ChatGPT-style window for prompts, with separate dropdowns for websites, images, and a dozen other visual formats. Once you put in the request, format, and style, you’ll be presented with a series of “A vs. B” choices until you’ve ranked the handful of outputs from best to worst.
It’s a useful service, but the real value of the platform comes from the enterprise side, where participating models can treat it as a source of endless instant feedback for their media-generating models. The users tend to be indifferent to which models they’re ranking — as Li puts it, they just want the best output they can get — so their rankings can give critical input to what users really want.
For frontier labs, that’s a service worth paying for, Li says, adding the site is currently generating $60 million in ARR, solidifying its position as a key source of human-led evaluation data for the AI industry.
Crucially, users have to log in to get their output, so Intelligence can also track how those tastes change across different continents and over time. (Li notes that web dashboards in Asia tend to have a more maximalist design style.) These measures are an important complement to automated benchmarks, which can operate at a greater scale but are often subject to being gamed or otherwise manipulated, as the Hugging Face breach demonstrated in dramatic fashion last week.
That’s not to say that crowdsourced human feedback will be an automatic winning market. Less than a year after launching, Yupp shuttered its doors earlier this year after raising $33 million from a16z crypto’s Chris Dixon. It too nabbed some frontier models as customers and had, it said, over 1.3 million users, but still couldn’t build a sustainable long-term business.
Even so, other startups based on human evaluation seem to be thriving. LM Arena, which takes a similar approach to text-based responses, raised $150 million in a Series A in January, just four months after formally launching its paid product.
Transform Any Place with Nano Banana in Google Earth
Google Earth briefly launched an AI-powered image generation feature that allowed users to visualize historical or speculative landscapes, only to roll it back hours later.
Summary
Decoder
- Nano Banana: The internal code name for Google’s generative model integrated into Earth for contextual scene reconstruction.
Original Article
Transform any place with Nano Banana in Google Earth
Update, 7/31: We know that people uniquely trust Google Earth for a reliable view of the world. We’ve seen geospatial professionals using this feature for a range of useful purposes, however we’ve also seen people sharing screenshots of generated imagery that appear to violate our policies. So we’re rolling back this feature in Google Earth while we work on implementing stronger guardrails. It's important to note that generated images didn’t appear in the main Google Earth experience for others to see and were watermarked as AI generated.
Original post, 7/30: Have you ever looked at an empty lot in your neighborhood and imagined a community garden, or wondered what your city looked like a century ago? Now Nano Banana 2’s image generation capabilities in Google Earth can show you both — and so much more.
For the first time, you can generate custom images using Google Earth’s satellite, aerial, and 3D imagery alongside Nano Banana, which creates concepts grounded in the real world. Just zoom in to a place in Google Earth on web, tap “create image,” and type whatever you want to see.
Here are five ways to try it.
1. Bring history to life.
Help students visualize the past instead of just reading about it. Teachers can show their classroom a historic site and type, “Render a hyper-realistic view of what these Pompeii ruins looked like in 78 A.D.” The modern-day ruins will instantly transform into a bustling, colorful street scene from the Roman Empire.
2. Learn more while you explore.
Create custom infographics to learn about a place without leaving Google Earth. Type, “Create an easy to understand infographic of the Statue of Liberty with key historical facts.” Behind the scenes, Gemini will retrieve relevant historical information and Nano Banana will create a graphic in seconds.
3. Create professional real estate plans.
Pitching a real estate project is much easier when your clients can visualize the final result. Architects and urban planners can now bring their plans to life by saying, "Reimagine this empty lot in Tokyo as a vibrant shopping and retail district with open spaces.” Barren concrete will be replaced by a high-quality 3D rendering of a bright urban oasis, helping your clients see what’s possible.
4. Visualize projects before breaking ground.
Whether you're dreaming up a new backyard studio or your lakeside dream home, it can be hard to picture the final result. Now, you can visualize it quickly. Zoom in to an open lot and say, "Add a modern lakefront cabin built from sustainably sourced local materials” to see a photorealistic rendering of your future house nestled perfectly in the actual landscape.
5. Give your favorite place a makeover.
Sometimes it’s fun to let your imagination run wild. To see what a place might look like a hundred years from now, type “Transform the Google Mountain View campus into a futuristic sci-fi utopia with glowing walkways, glass biodomes, flying transport pods, and trees wrapped around buildings,” and watch the area morph into a vibrant, cyberpunk-inspired metropolis right out of a movie.
Image generation with Nano Banana is available globally today for Google Earth web users. Head over to Google Earth, pick a spot on the map, and start bringing your ideas to life.
New Lenovo leak just gave us our best look at Google's post-Chromebook gamble
Leaked images of the upcoming Lenovo Googlebook 15 suggest Google is attempting to pivot Android into a viable desktop operating system for laptops.
Summary
Original Article
Leaked images of the Lenovo Googlebook 15 reveal Google's upcoming Android-powered laptop, which features a premium design, a dedicated Gemini key, a desktop-style Android interface, and a wide selection of ports aimed at competing directly with Windows laptops. The new interface combines a phone-like status bar with a traditional taskbar, while preloaded apps such as Chrome, Gmail, Canva, and Adobe Premiere Pro highlight Google's ambition to make Android a serious desktop platform. Although specifications, pricing, and performance remain unknown, the laptop is expected to launch this fall, potentially at IFA in September.
Creative Technologists are moving from the margins to the center
The role of the Creative Technologist is transitioning from an obscure advertising niche to a central function in product teams as AI changes how we build.
Summary
Deep Dive
- Historical Evolution: The role evolved from 1960s experimental art labs like E.A.T. to modern product-focused design engineering.
- Defining the Role: It sits at the intersection of experimentation, strategy, and implementation.
- The Impact of AI: AI is accelerating the need for people who can explore new capabilities and translate them into meaningful user experiences.
- Organizational Value: Companies are poaching these individuals because they provide a competitive edge in defining original use cases for new, unstable technologies.
- Future Outlook: The mindset of the 'Creative Technologist' is becoming a universal requirement for modern product designers, regardless of job title.
Decoder
- Creative Technologist: A professional who uses technology as a medium for creative expression and rapid prototyping to discover new product applications.
- Experiential Studio: Agencies focused on building interactive, physical, or immersive installations for brands.
Original Article
Creative Technologists are moving from the margins to the center
Once an obscure job title, creative technology is becoming a new way of working, and a core capability inside modern product teams. Soon, many of us may be working this way.
As AI collapses the distance between idea and implementation, companies are looking for people who can move across design, code, systems, and storytelling.
Let us begin by disambiguating the Creative Technologist role. Everyone references it as though we all know what it means. But when I ask people to define it — with specificity — most cannot.
And it’s understandable, the role is legitimately hard to pin down. It has existed for decades, mostly in advertising agencies, research labs, and experiential studios. It has also gone by many names: UX engineer, design technologist, creative developer, UX prototyper, and design engineer. The taxonomy is a mess. Even people who hold the title admit it can sound entirely made up.
But the Creative Technologist is having a renaissance. The role is being pulled from the artistic margins into the center of the machine, and understanding its origin story and direction tells us something about where many creative careers are going.
So let me try to disambiguate it.
Instead of beginning with definitions, let me start with a person.
On a recent Google Flow creator community call, the guest was Tina Tarighian, a Senior Creative Technologist at Google Creative Lab in New York. Her first-person essay, “WTF is a Creative Technologist?”, published by Will Zimmerman’s 3rd Space, is one of the clearest accounts of the role I have read. If you read one other piece on this subject, can it be that one.
Her path: a nationally ranked competitive debater in high school. A political science major who thought she might become the next Hillary Clinton. A computer science degree that made her “actually study, sob, and then study some more.” A hedge-fund internship in Tokyo that felt like wearing someone else’s shoes. A psychic who pulled her off the street to tell her she was in the wrong career.
Then a net artist in Greenpoint taught her JavaScript. She was broke enough to walk 45 minutes to his studio rather than take the train. She began cold-pitching companies with creative technology ideas. Her first paid project was a virtual try-on filter for a compression-underwear brand. She was paid almost nothing and remained, in her words, “insufferably giddy for weeks.”
She cold-messaged a musician the day before he started a band and ended up live-coding visuals at its debut show. One project turned into ten: a New Museum residency, fashion-show installations, gallery work, and her own runway shows. When she missed her no-contact ex-girlfriend, she built an email lottery with a tiny chance of actually sending her note. Five million people saw it, and people in more than 100 countries used it.
Now she works at Google Creative Lab, demonstrating tools built over a weekend to arenas of thousands, collaborating with Paris Hilton on ADHD-friendly apps using Gemini, and presenting on the main stage at Google I/O.
Here is how she describes the actual work:
“We are writing the first draft of what a technology means, sitting with something new before the obvious use cases have had a chance to calcify, and trying to find something meaningful in it.”
That is the role: writing the first draft of what a technology means.
The two ends of the spectrum
Tina maps the field better than any formal taxonomy I have found.
On one end is the commercial version, shaped by the dot-com boom, when companies needed people who could “talk to people and computers with the same suavity.” Some lean toward hardware, some live in the browser, and some build strange physical objects that end up in advertising campaigns. The through-line, in her words, is that “they’re technical enough to build and perspicacious enough to know what’s worth building.”
On the other end are people using technology primarily as an artistic medium. Think of Refik Anadol turning data into large-scale immersive environments, generative systems that feel alive, or Tina’s email lottery: code used as emotional expression. “What we all have in common,” she writes, “is that we are using technology as our primary medium to say something, not to ship products.”
Blair Neal, a veteran of the field who spent years at the experiential studio Fake Love, has been documenting this territory for more than a decade. His “Advice for Creative Technologists” is canonical field reading, and his writing on creative technology inside organizations offers the explanation I would give my parents:
“I am an artist, and just like a painter’s medium is the full range of paint types and canvases, my medium is technology itself.”
Neal also created A Creative Technology Taxonomy, a large map of what a senior creative technologist might need to understand, from projection mapping and sensors to game engines and machine learning. Not everything needs to be mastered. The value of the map is that it provides a knowledge foundation for making informed choices about which technologies serve which experiences.
So when someone asks what a Creative Technologist is, the sincere answer is: someone whose medium is technology itself. Sometimes that medium is applied to commerce, sometimes to art, and often to an in-between space without a clean name.
A brief history: 60 years of the same person with different names
The thing that gets lost in the “hot new role” discourse is that this archetype is not new. Its origin story goes back six decades, and tracing it reveals something important about where it is heading.
1966: the founding moment. Bell Labs engineers Billy Klüver and Fred Waldhauer teamed up with artists Robert Rauschenberg and Robert Whitman to stage 9 Evenings: Theatre and Engineering at the 69th Regiment Armory in New York. Ten artists, including John Cage and Yvonne Rainer, worked with more than 30 Bell Labs engineers and scientists. The performances used closed-circuit television, infrared cameras, Doppler sonar, and other emerging technologies. The collaboration led to Experiments in Art and Technology, a nonprofit built around pairing artists and engineers. This is one of the clearest institutional origin points for creative technology.
1973: the academy takes over. Muriel Cooper co-founded the Visible Language Workshop at MIT, exploring how computation could transform graphic design, typography, and publishing. The workshop later became one of the founding research groups of the MIT Media Lab. The question moved from “Can artists use technology?” to “Is computation itself a design medium?”
1996–2001: the medium gets its language. John Maeda led the Aesthetics + Computation Group at the Media Lab and argued that designers should understand the computer as a medium, not just a tool. His Design by Numbers project taught programming to visual artists. Two of his students, Casey Reas and Ben Fry, later created Processing, a digital sketchbook where designers and artists could code. As documented in AIGA Eye on Design’s oral history of Processing, it became both a software environment and a worldwide creative-coding community. In parallel, the dot-com boom produced the commercial variant: companies needed people who could move between technical infrastructure and creative intent.
2005–2015: the agency golden age. Arduino made hardware experimentation more accessible. openFrameworks and Cinder gave creative coders professional tools. Flash turned the web into an expressive playground. Agencies such as R/GA and Wieden+Kennedy, alongside studios such as Fake Love, helped build the experiential-marketing industry around creative technologists. Interactive installations, projection mapping, and sensor-driven brand experiences made “Creative Technologist” a role people could apply for, even when few people outside the field understood it.
2015–2022: fragmentation. The role splintered as product companies absorbed parts of it. Amazon used Design Technologist. Google used UX Engineer. Startups used Creative Developer. Product-focused branches moved toward what the industry now frequently calls Design Engineer. The artistic branch continued independently, from large-scale data sculpture to small experimental internet communities.
2022–now: the AI renaissance. Companies suddenly hold technologies whose cultural and experiential meaning has not yet been written. The capability that Experiments in Art and Technology cultivated in 1966 — sitting with a new technology before its use cases calcify and finding something meaningful in it — has become scarce again.
The pattern across these decades is consistent: every time a genuinely new technology arrives — video, computation, the web, physical computing, AI — the people who figure out what it means are often the ones who treat it as a medium rather than just a tool. The names change. The disposition does not.
How it maps against adjacent roles
Where does this sit relative to the other roles emerging around design and technology?
Here is my attempt at an ontology.
Creative Technologist: the broadest and most experimental role. It combines technology, creativity, and strategy across hardware and software, installations and browsers, art and advertising. It is comfortable with the question, “What could this technology mean?” before anyone knows what to build. Google Creative Lab is one archetype: its roles have described work translating early-stage technologies into tangible, engaging experiences.
Design Technologist: the product-focused version. Companies such as Amazon and eBay have used this title for roles involving heavy prototyping, design systems, and UX engineering. It is generally more structured and less open-ended.
Design Engineer: the most production-oriented version. It combines product design, front-end code, and the craft of shipping interfaces. Companies such as Vercel, Stripe, Cursor, and Lovable have helped make the title increasingly legible in product teams.
The boundaries are porous, but the orientation differs. The Creative Technologist asks, “What could this mean?” The Design Engineer asks, “How do we ship this well?” One writes the first draft of a technology’s meaning; the other makes sure the final draft is crafted.
Why this role is having a moment
Three things are converging.
First: AI needs first drafts. Every company suddenly has access to technologies whose obvious use cases haven’t calcified yet. Models drop monthly. Capabilities shift weekly. And the question inside every organization is exactly the one Creative Technologists are trained to answer: what does this mean for us? Tina describes the actual brief at Google: a new model drops internally, and the ask is some version of “help people feel something other than dread about this”. That is the job now, at every company.
Second, AI has lowered parts of the technical barrier, democratizing the technical layer. The earlier Creative Technologist often needed Processing, openFrameworks, C++, custom shaders, and hardware expertise. Those barriers kept the role relatively niche. A new generation can prototype with coding agents and generative models, although expertise still determines whether those prototypes become robust, responsible experiences. Adobe’s recent report, “The evolving role of the Creative Technologist,” describes the role as connective tissue across creative, engineering, marketing, and other teams — and as a missing layer in many enterprises.
Third: the doing is being celebrated. Tina again: “My fellow CTs are being flown out to give talks, poached mid-job with huge pay bumps, and inundated with DMs from startups begging them to find a use for their product.” The people who were building weird internet art for years are suddenly the most in-demand profile in tech. Product teams at Google. Design systems at Apple. Experimental storytelling at A24 in partnership with DeepMind.
Will we all become Creative Technologists?
Here is the question behind the question, and my current answer: not necessarily in title, but increasingly in how we relate to technology.
Tina makes an observation that stays with me:
“Today, kids are building entire worlds in Roblox like it’s a normal extension of playing. People are making little tools for their friends, their classrooms, or their own weird niche problems. No one waits for permission or funding or someone to translate their idea into code. Coding just feels like… a thing people do now.”
Her conclusion is that “Creative Technologist” may become less of a mutant profession that requires explanation and more of a way of relating to the most pervasive medium on earth.
That’s the real thesis. The Creative Technologist was never really a job title, it was an early name for a relationship to technology that’s now becoming universal: technology not as an industry you work in, but as a medium you think with.
What this means for you
You don’t need to change your title. You need to notice whether you’re relating to technology as an industry or as a medium.
The industry relationship: you follow the tools, learn what you must, wait for the use cases to be defined, and then execute within them.
The medium relationship: you sit with new technology before its meaning has crystallised. You build small things to understand it viscerally. You ask what it could mean, not just what it does. You use it to say something.
The second relationship is what Tina found through nightclubs and net art. It’s what Blair Neal describes as choosing technologies “in thoughtful ways that honor the person experiencing something”. It’s what I keep trying to practice with every experiment I run and every “unsolicited” prototype I build.
It is learnable. It can start with one small project in which you use technology to say something, instead of simply shipping something.
That is the whole practice. The title is optional.
further reading
- WTF is a Creative Technologist? — Tina Tarighian, 3rd Space
- Advice for Creative Technologists — Blair Neal
- A Creative Technology Taxonomy — Blair Neal
- Creative Technology and Organizational Structures — Blair Neal
- The Evolving Role of the Creative Technologist — Adobe
- How to Become a Creative Technologist in 2026 — Ayca Turan
- What the Heck is a Creative Technologist? — John Evanofski
- The Best Paid Emerging Creative Roles in 2026 — Creativepool
- aesthetic.computer — Jeffery Alan Scudder
- Experiments in Art and Technology (E.A.T.)
- Processing: the Software that Shaped Creative Coding — AIGA Eye on Design
10 GUI Design Elements Build Every User Interface
Jakob Nielsen defines the 10 core interface elements that have served as the foundation of user interface design for over four decades.
Summary
Decoder
- GUI (Graphical User Interface): A visual way of interacting with a computer using items like icons, menus, and windows.
Original Article
Nielsen identifies 10 fundamental GUI elements—buttons, forms, menus, links, dialogs, alerts, icons, checkboxes/radio buttons, tabs, and search—that form the basic vocabulary underlying nearly every interface built over 40+ years. The article traces each element's history, explains its usability role, and offers 86 evidence-based design guidelines plus bonus coverage of windows and pointers. It argues designers should master this established alphabet rather than invent new conventions, since familiarity, tested through Jakob's Law and decades of use, consistently outperforms novelty.
Measuring the UX of AI
Measuring generative AI UX requires shifting from traditional usability testing to frameworks that account for model variability and user trust.
Summary
Deep Dive
- Traditional usability metrics like 'time on task' are less meaningful when the AI's processing time fluctuates significantly.
- 'Interaction success' must be redefined to account for the truthfulness and relevance of generative responses, not just task completion.
- Measuring user trust is critical because a chatbot that provides incorrect information with high confidence can negatively impact product perception.
- Behavioral data, such as prompt refinement counts, provides insight into how users struggle with system ambiguity.
- Attitudinal surveys should specifically probe the user's perception of AI capability compared to their initial expectations.
Decoder
- SUS (System Usability Scale): A 10-item questionnaire used to provide a global view of subjective assessments of usability.
- Non-deterministic: A property of a system where the same input does not guarantee the same output, common in stochastic LLM generation.
Original Article
Generative AI chatbots' UX can be measured through a mix of behavioral and attitudinal metrics.
SpaceX's Bumpy Ride on the Stock Market May Get Bumpier
SpaceX employees and early investors can finally trade shares today as the company's initial post-IPO lockup period expires.
Summary
Original Article
The first lockup on SpaceX shares is set to expire today. Employees and other company insiders who were prevented from trading stock after SpaceX's IPO will now be able to do so. Lockups prevent shares from flooding the market and driving down their price after an IPO. More than double the current supply that can be traded will be made available, which may further depress the price.
The Hidden Details in Every Font
Typography succeeds when it is invisible, as designers hide centuries of visual adjustments behind the text to ensure effortless reading.
Summary
Decoder
- Overshoot: The extension of rounded letters beyond the baseline to make them appear equal in size to flat-bottomed letters.
- Ink Trap: Small gaps cut into the corners of letterforms to prevent ink from pooling and blurring the shape when printed at small sizes.
Original Article
Good typography is largely invisible. Type designers spend months making subtle adjustments—such as optical sizing, spacing, overshoots, ink traps, and character differentiation—so that text feels effortless to read without readers ever noticing the craft behind it. Many of these techniques have been refined over centuries and are based on how human vision works, ensuring letters remain legible at different sizes, in different contexts, and even under challenging conditions like poor printing or road signage. The result is that typography doesn't just convey words - it quietly influences readability, tone, and even credibility, demonstrating that the best type design succeeds precisely because it disappears behind the meaning.
Turn Data Into Winning Ads (Website)
Adomate allows users to build custom AI workflows for converting performance data and consumer signals into ad creatives.
Summary
Original Article
Turn performance data, competitor insights, and consumer signals into winning ad creatives, using custom AI workflows you build and own.
Designer's attempt to fix Disney+ logo fires up fans
A designer's attempt to 'fix' the Disney+ logo sparked fan backlash, highlighting the extreme cultural sensitivity of legacy brand marks.
Summary
Original Article
Designer Allan Peters proposed a simplified Disney+ logo that uses the iconic Disney "D" as the main mark, highlighting a hidden "+" already present within the letter. While many appreciated the clever concept, the redesign sparked backlash from fans who argued that Disney's signature "D"—closely associated with Walt Disney's handwriting—is too iconic and culturally significant to alter, illustrating the challenge of updating one of the world's most recognizable brand identities.
Koto on why physical space is still critical to creative culture
Creative agency Koto argues that physical studio space is non-negotiable for fostering mentorship, high-craft standards, and spontaneous collaboration.
Summary
Original Article
Koto believes a strong in-person studio culture is essential for creative work, arguing that being together enables spontaneous collaboration, mentorship, relationship-building, and higher creative standards in ways remote work cannot. The agency intentionally designs its offices around collaboration rather than capacity, while allowing each studio to develop its own personality through local culture, team-led rituals, and thoughtfully designed spaces. As Koto grows, it prioritizes hiring ambitious, low-ego people with potential, investing in shared experiences, and preserving a culture of optimism, craftsmanship, and co-creation to build a lasting creative legacy.
Everyone's Work is Starting to Look the Same
Design homogenization driven by generative AI is empirically measurable but currently limited to highly constrained tasks like user onboarding.
Summary
Deep Dive
- AI tools tend to suggest the most statistically probable patterns, which often align with established 'best practices.'
- Research shows the effect is most pronounced in UI tasks with fixed, objective goals.
- Designers are shifting from creation to curation of AI-generated assets, increasing the likelihood of aesthetic convergence.
- The pressure to ship quickly using template-heavy AI workflows exacerbates the issue of 'sameness.'
- Over-reliance on generative outputs may diminish the long-term development of original, brand-specific interaction design.
Original Article
AI is homogenizing design output, but peer-reviewed research shows the effect is measurable yet small, strongest in tightly constrained tasks like onboarding flows.