An undercover Google analyst infiltrated a notorious supply-chain hacking gang
A Google threat analyst acted as a 'fly on the wall' inside TeamPCP, a prolific supply-chain hacking group that compromised over a thousand companies.
Summary
Deep Dive
- TeamPCP automated supply-chain attacks by tainting widely used open-source libraries.
- The group's targets included major infrastructure and security tools, cascading into breaches at companies like OpenAI and GitHub.
- The infiltration allowed Google to download and patch a zero-day exploit the group was developing using AI.
- In-fighting between cybercriminal groups (specifically a betrayal by ShinyHunters) helped accelerate the collapse of the group's operations.
- Google identified the perpetrators through poor operational security, specifically the use of a personal Google Drive account to store stolen credentials.
Decoder
- Supply-chain hacking: The act of compromising a software component (like an open-source library) used by others, allowing the attacker to infect the downstream products that rely on that component.
- Zero-day exploit: A cyberattack that leverages a software vulnerability unknown to the software developer, meaning there is no existing patch.
Original Article
Before two of its alleged members were arrested and charged in Australia last month, the hacker group known as TeamPCP carried out a hacking spree unlike any other in history. It tainted hundreds of open-source programs with its malware, stole developer accounts to perpetuate that software supply-chain hacking, and even released a Dune-themed self-spreading worm to automate the process, ultimately breaching more than a thousand companies.
Now Google’s threat intelligence group has revealed that during a key moment of TeamPCP’s rampage, the company’s own undercover researcher had infiltrated the group—allowing Google to monitor the hacking spree from the inside, warn breach targets, and even help disrupt the group’s attempts to exploit those victims.
In a talk at security firm SentinelOne’s LABScon research conference today, Google Threat Intelligence Group researcher Austin Larsen will present details on the company’s investigation—and infiltration—of TeamPCP amidst the group’s unprecedented, chaotic supply-chain hacking campaign. According to Larsen, Google eventually followed a trail of operational security mistakes allegedly made by one of the two Australians now accused of being leading members of the hacker group and passed on key identifying details to law enforcement. The company also received intelligence from ShinyHunters, another infamous cybercriminal group that TeamPCP partnered with, but which later turned on the supply-chain hackers. And perhaps most surprisingly, Larsen says that Google’s security subsidiary Mandiant had an undercover analyst—not himself—within the group’s inner circle from almost the beginning of TeamPCP’s time in the spotlight.
“One of our personas had been working for many months to build trust with one of the actors that was invited to join TeamPCP, and so was added to the group,” Larsen told WIRED in an interview ahead of his LABScon talk. “So essentially, almost day one, Mandiant was watching everything behind the scenes.”
The TeamPCP mole
Late last month, Ruben Ian Thomson and Louis Michael Gaebler, both Australians in their early 20s, were arrested by Australian police in a joint investigation with assistance from the FBI, charged with hacking crimes, and described by the Australian Federal Police (AFP)—in a press release that, due to Australian privacy laws, did not name them—as “principal participants” in TeamPCP. The hacker group, which seems to have first appeared online in late 2025, had made headlines with a brazen string of cascading supply-chain attacks: It repeatedly compromised open-source software to hide its malware, which then allowed it to hijack the credentials of software developers and plant its malicious code in yet another widely used tool, in a repeating cycle.
Starting this spring, for instance, TeamPCP compromised the open-source security scanner Trivy, the AI application programming interface tool LiteLLM, infrastructure of the web application security firm Checkmarx, the web app library TanStack, and the enterprise AI platform Mistral AI. Those repeated supply-chain attacks, with each enabling the group to cast its net again for more victims, ultimately allowed the hackers to breach open-source code repository GitHub, data contracting firm Mercor, and employee devices at OpenAI, the European Commission, and many others who have remained unnamed in public reporting. At times, the group deployed a worm known as Mini Shai-Hulud, named after the sandworms in Dune, to automate its hacking and scale up to even more victims. (The name also seemed to refer to an earlier Shai-Hulud worm that hackers designed to try a similar approach in September 2025, though it’s still not clear if TeamPCP or any of its alleged members were involved in that earlier intrusion campaign.)
Larsen now says that in March, just as TeamPCP was beginning its frenzied supply-chain hacking, Google’s own undercover analyst was invited to join the hackers’ inner circle. That inside source, whose name Larsen declined to reveal, was one of about 12 members of the group given access to a core chat that TeamPCP called CanisterWorm.
“You guys should understand that we pulled off the biggest supplychain [sic] maybe ever recorded in modern history,” one TeamPCP member wrote in the leaked chats.
Michael Fletcher, a former AFP analyst who now works in the threat research division of an Australian telecom firm, says he approached Larsen around that time about methods for monitoring the group’s members and activities. He says that Larsen responded by asking Fletcher to approach the hackers with caution because one of them was a “friendly,” Fletcher remembers. “I thought, damn, you all have been inside this early,” he says.
Google’s undercover analyst, Larsen says, gained access to a server where TeamPCP was storing its trove of credentials stolen from its many victims: the usernames, passwords, and access tokens it had obtained through its hacking and seemingly planned to use to extort target companies. So Google’s team decided to take action to warn victims and prevent TeamPCP’s ransom scheme. “My thought was: How can we, as quickly as possible, disrupt their campaign before more compromises can happen?” Larsen says. “Let’s go mess up what they’re doing. That was my goal.”
Rather than focus on alerting the owners of the stolen credentials at victim companies directly, which Larsen says would have taken too long given the sheer number of breached companies, Google first reached out to providers where those credentials could be used, like Amazon Web Services and Microsoft, to have the credentials revoked and prevent the hackers from exploiting them. Larsen and his team sent out hundreds of notification emails to those providers and then to victims, many of which got immediate responses.
Around the same time, Larsen says, Google’s visibility into the TeamPCP internal chat also allowed it to learn that someone within the group’s core circle was, distinct from the group’s supply-chain hacking, using an AI tool to develop a zero-day exploit in a widely used piece of login software that would allow the hackers to bypass its two-factor authentication. Google’s team got a copy of the exploit code, tested it out, and found that, with a few tweaks, it worked—a rare instance of an in-the-wild AI-created hacking technique that took advantage of a previously unknown software vulnerability. Google warned the software’s developer, who was able to patch its security flaw. (The incident was described in a case study Google released in May, but without naming TeamPCP or detailing how Google learned about the exploit.)
More betrayals, sloppy opsec
Google’s analyst was not, it turns out, the only traitor in TeamPCP’s midst.
Even prior to Google’s disruption effort, Larsen says, the group struggled to profit from its enormous collection of stolen data, which, according to the AFP, included more than half a million users’ credentials. Larsen estimates that, despite that haul, it was pulling in only tens of thousands of dollars in extortion payments, not the millions similar groups have amassed. So in an attempt to better monetize its hacking, TeamPCP invited multiple other cybercriminal groups to partner with it, giving them access to the stolen credentials in exchange for a percentage of any extortion payments they were able to extract.
One of those cybercriminal partners was ShinyHunters, a years-old, highly prolific hacker group that has extorted millions of dollars from victims through data theft and ransomware, including in the breach of educational software platform Canvas that would later paralyze thousands of schools across the US. Around April, a few weeks after partnering with TeamPCP, ShinyHunters went rogue, Larsen says, carrying out its own extortions with TeamPCP’s credentials but without giving the supply-chain hackers their cut. ShinyHunters went so far as to share with Larsen, unsolicited, a full log of the group’s chat on TeamPCP’s server—not knowing that he already had access via Google’s mole.
ShinyHunters also taunted TeamPCP in messages on X, and its louder betrayal got the latter group’s attention. TeamPCP responded by narrowing its inner circle, moving its data to a new server, and exiling ShinyHunters and several other group members from its CanisterWorm chat, including Google’s undercover analyst.
“Just delete that and stop sharing shit with shinyhunters,” one of the TeamPCP leaders wrote.
Even without that inside source, though, Larsen says more traditional digital detective work allowed him to piece together the trail of breadcrumbs that would ultimately let him learn the identity of Thomson, one of the two men charged for allegedly playing “key” roles in TeamPCP. Larsen found in a leak of user data from the BreachForums hacker forum that one of the most active handles in the CanisterWorm chat had been registered with the Gmail address sheepstealing@gmail.com. Combing through other forum archives, he found a 2019 dispute between someone with the pseudonym sheepstealing and a seller of pirated Microsoft Office keys, in which the sheepstealing user demanded a refund at a PayPal account tied to the email ruben@thomsonfamily.net.au.
After TeamPCP moved its stolen credentials to a server hosted by a different provider, Larsen says, Google was able to learn about some contents of the new server—through what Larsen describes as a “trusted partner”—and also that it was being backed up to a Google Drive on that same sheepstealing@gmail.com account.
“When we saw that, I just thought: There’s no way. Why would he be sending all of this illicit, stolen material to a Google Drive that’s tied to himself?” Larsen says. “That’s when we gave the tip to the FBI.” Larsen says he got an interested response from an agent in a matter of minutes.
In a statement to WIRED, the FBI declined to comment on any “active investigation” but noted that it “is able to confirm we strive to increase impact on adversaries through partnerships as documented in our newly released FBI Cyber Strategy.” The AFP declined to comment.
About a month after his tip, Larsen says, US law enforcement had finished the legal process of requesting Thomson’s data from Google with a warrant. Late last month, Thomson was arrested by Australian police, who released a video of him being walked out of a suburban home in a Northface hoodie and sweatpants.
Neither Thomson nor Gaebler, the other alleged member of TeamPCP who was arrested, could be reached for comment.
Larsen was careful to note that Google’s undercover analyst inside of TeamPCP never engaged in any illegal hacking or even encouragement of the group’s breaches. “They were a fly on the wall, only saying enough to not be suspicious,” Larsen says. “There are guardrails around what we do.”
But Larsen also notes that his team’s work to actively foil TeamPCP’s hacking is part of a new shift within Google. The investigation, after all, kicked off around the same time as Google’s newly launched Cyber Disruption Unit, which has been officially tasked with taking a more aggressive approach to combating cybercrime and state-sponsored hacking.
“Google Threat Intelligence Group has put an emphasis on disruption. That’s one of our missions now,” Larsen says. “Writing reports can only be so useful. Taking action to protect users and customers—that is the next step.”
Frontier Overhangs
Ben Thompson argues that calls to 'pace' the AI frontier are less about safety and more about managing business risks and competitive moats.
Summary
Deep Dive
- The rise of the agentic paradigm (notably since Opus 4.5) has validated the need for massive compute infrastructure.
- Model-harness integration is becoming the primary source of differentiation, threatening to commoditize model providers.
- Microsoft’s Copilot strategy, which allows users to swap models within a consistent harness, signals a shift toward modularity in the AI stack.
- Data retention policies (e.g., Anthropic's Fable 5.1 changes) serve as evidence that enterprise customers will prioritize business requirements over pure frontier model performance.
- Meta’s Muse indicates that 'good enough' model performance can still build highly valuable, sticky consumer products.
- Frontier labs are under pressure to own user touchpoints to survive, putting them on a collision course with traditional software companies.
- Cyber defense requires full automation to combat offensive AI agents, which paradoxically requires pushing the frontier faster rather than slowing down.
- 'Safety' arguments are often used as strategic cover by incumbents to slow competition while they build defensive moats.
Decoder
- Harness: The software infrastructure that connects an AI model to external tools, APIs, and user data to perform tasks.
- Overhang: A market condition where demand significantly exceeds supply (e.g., compute availability) or prices are artificially inflated by scarcity.
- Modular architecture: A system design where independent components can be upgraded or swapped without redesigning the entire system.
- Agentic paradigm: The shift from using AI for simple text generation to using AI to execute complex, multi-step actions and tool-use autonomously.
Original Article
Frontier Overhangs
There has been, over the last week, what I think is a healthy debate about the philosophy and psychology that undergirds the views of meaningful segments of the AI community, particularly those obsessed with doomsday scenarios. It is, in the end, difficult to reason with a philosophy that grants equivalent moral weight to not just all beings — human or not — who exist today, but who may ever exist in the future; this tilts the scales in such an absurd fashion towards safetyism that innovation is impossible and freedom is intolerable.
Worse, it taps into the psychology of religion, where dissent is not brooked and questioning the premise is heresy. I guess that makes me a heretic then: I reject the premise in favor of doubt in our ability to foresee the future, combined with faith in humanity figuring things out along the way. If that leads to the most fantastical doomsday scenario, then I think the State of New Hampshire said it best:
I’m referring, of course, to the question of Effective Altruism and the extent to which it is intermingled with Anthropic and CEO Dario Amodei’s insistence that We Must Pace the Frontier. I expanded on my objections to some of the philosophical implications in the last | two episodes of Sharp Tech, but Stratechery is a site about strategy and technology, and those angles deserve examination as well.
Three months ago I explained in Anthropic’s Safety Superpower how the company’s genuine belief in its safety rhetoric conveniently gave it license to pursue extraordinary goals that aligned with its business interests, including aggressive attempts to disintermediate software, collect sacrosanct customer data, and secretly sabotage would-be competitors. This analysis is in a similar vein but from the opposite direction: “Pacing the Frontier” is framed as — and I believe motivated by — a concern about safety, but it also happens to address several distinct problems faced by the frontier labs. These problems take the form of overhangs that have formed by virtue of how rapidly models are improving; slowing model improvement would reduce the overhangs.
The Capability Overhang
It was six months ago that I wrote Agents Over Bubbles, where I argued that the agentic paradigm, which kicked off with the release of Opus 4.5 in late November 2025, was so incredibly capable and so incredibly token hungry that we were not in an infrastructure bubble: we really do need all of the compute that is being built.
I do, as a matter of my job, talk about things that I do not necessarily experience personally; the canonical example is my long-running analysis of and advocacy for advertising as a business model, despite the fact I run a subscription business. I mention this because I developed this take on agentic AI before I dove headfirst into agentic coding for myself, and just as well: I have been so taken by the possibilities of making software for myself that I have started multiple new projects over the past few months, and I’ve even finished a few of them! It’s thrilling, and the possibilities seem endless. Needless to say, I believe this more than ever.
At the same time, there is one part of that Article that I’ve wavered on, and that is the importance of model company harnesses:
Specifically, I noted above that what made Opus 4.5 compelling was not the model release itself, but changes to the Claude Code harness that made it suddenly dramatically more useful. What this means is that model performance isn’t the only thing that matters: the integration between model and harness is where true agent differentiation is found.
This is a very big deal when it comes to figuring out the future structure of the AI industry and where profits will flow, because profits flow away from modular parts of the value chain — which are commoditized — and flow towards integrated parts of the value chain, which are differentiated…It follows, then, that if agents require integration between model and harness, that the companies building that integration — specifically Anthropic and OpenAI — are actually poised to be significantly more profitable than it might have seemed as recently as late last year. And, by the same token, companies who were betting on model commoditization may struggle to deliver competitive products.
One of the projects I’ve undertaken over the last month or so actually entailed — at least in one of its permutations — building my own harness, and it’s doable! It also is very difficult, at least with my level of agent-mediated capability; that noted, my go-to example in that Article about how much harness-model integration matters was Microsoft anchoring its new E7 enterprise offering around Claude Cowork, but CEO Satya Nadella broke the news in a Stratechery Interview that that was only a temporary state of affairs:
We’re using the same harness that we use in GitHub and the same thing in security, too. So we have the same harness that’s a multi-model harness in which we will rotate through — obviously MAI by default gets trained in our harness, but we will have GPT, we will have Anthropic in there and any open weight model. We will allow anyone to take any of the models they fine-tune or build. In fact, they can take an open weight model from Fireworks, tune it, put it into Copilot, no problem.
It took a while for Nadella’s claims to reflect shipping reality, but sure enough you can now choose a model for Copilot Cowork; it will be up to end users to determine how competitive Microsoft’s harness is with Claude’s, but clearly the harness and the model can be different things.
This is actually a more meaningful development than it might seem, for reasons that go back to the theory of integration and modularity put forward by the late Clayton Christensen; from The Innovator’s Solution:
The left side of figure 5-1 indicates that when there is a performance gap — when product functionality and reliability are not yet good enough to address the needs of customers in a given tier of the market — companies must compete by making the best possible products. In the race to do this, firms that build their products around proprietary, interdependent architectures enjoy an important competitive advantage against competitors whose product architectures are modular, because the standardization inherent in modularity takes too many degrees of design freedom away from engineers, and they cannot optimize performance…
Once their requirements for functionality and reliability have been met, customers begin to redefine what is not good enough. What becomes not good enough is that customers can’t get exactly what they want exactly when they need it, as conveniently as possible. Customers become willing to pay premium prices for improved performance along this new trajectory of innovation in speed, convenience, and customization. When this happens, we say that the basis of competition in a tier of the market has changed.
The pressure of competing along this new trajectory of improvement forces a gradual evolution in product architecture, as depicted in figure 5-1 — away from the interdependent, proprietary architectures that had the advantage in the not-good-enough era toward modular designs in the era of performance surplus. Modular architectures help companies to compete on the dimensions that matter in the lower-right portions of the disruption diagram. Companies can introduce new products faster because they can upgrade individual subsystems without having to redesign everything. Although standard interfaces invariably force compromise in system performance, firms have the slack to trade away some performance with these customers because functionality is more than good enough.
One of my go-to examples in Anthropic’s Safety Superpower was the company’s decision to predicate Fable usage on Anthropic holding onto all customer data for at least a month; this was a big deal, and I argued at the time that Anthropic was making a bet that its models were good enough to convince enterprises to give up on zero data retention:
It’s pretty significant, I think, that Anthropic is declaring that not retaining data is no longer an option, at least if you want access to their best models. Yes, today, that retention is for safety purposes only, and not for training; it’s plausible, however, that Anthropic’s lead becomes so significant that they quietly announce that they are going to train on that data as well, and companies will feel they have no choice but to go along. That additional training data, of course, will only further increase Anthropic’s lead, and all of this will be justified because Anthropic has already clearly decided they are the only ones who can be trusted to be in charge.
In fact, Fable wasn’t good enough: customers pushed back, and Fable usage stayed relatively low; when Fable 5.1 was released, the Anthropic-gets-to-keep-your-data provision was gone. This is evidence of Christensen’s theory in action: customers demonstrated the willingness to base their model-choice decision on something other than pure performance, namely, data retention policies.
This doesn’t, in and of itself, suggest that new model capabilities aren’t desired; it does, however, suggest that current model capabilities are “good enough” for customers to not do whatever is necessary to get access to the cutting edge, which reduces the value of the cutting edge to its proprietors, and gives credence to the strategy of Microsoft and others focused on separating harness and model. Pure capability no longer translates directly into a moat.
The Product Overhang
The modularization of models and harness explains why the frontier labs have what I called an economic imperative to own end user touchpoints; again from Anthropic’s Safety Superpower:
It has long been clear to me that the frontier labs have the economic imperative to move closer to the user. If you own the user touchpoint, then you have meaningful lock-in, and the best way to own the user touchpoint is to be the canvas for everything they need to do. This, by extension, means that the frontier labs are on a collision course with software companies: it’s software that owns the user touchpoint, and it’s in the frontier labs’ long-term interest to not simply be a commodity input into software but to simply replace software outright.
The key for the frontier labs, then, is to build those user touchpoints while they have superior capabilities. However, this is where Meta’s recent launch of Muse is a bearish signal. Muse is, by a significant margin, the best and most approachable personal agent product I have tried. Meta deserves a tremendous amount of credit for the product work they have put in, as well as the massive infrastructure commitment entailed in providing users with a very capable virtual machine for free. Oh, and of course they deserve credit for the Muse Spark model undergirding Muse.
That noted, Muse Spark 1.3, the most advanced Meta model, is still not state-of-the-art, and that is the bearish signal: it is good enough for a very good personal agent product, and critically, a personal agent is much stickier than a chatbot. Once you have put all of your information into a personal agent and actually incorporated it into your day-to-day life, it is much more of a challenge to change to something else. This is in contrast to Codex/ChatGPT and Claude Code: yes, you may have developed your own set of skills and understanding of how each harness works, but at the end of the day the relevant artifacts (i.e. your code) are in GitHub, and it’s not that much of a lift to point a different agent and harness at those artifacts if the alternative is better and/or cheaper (or doesn’t want to keep all of your data).
In short, model capability is good enough that compelling products — products that actually have moats — can now be built, and from a business perspective it would do Anthropic and OpenAI good to devote more of their resources to actually building such products.
The Pricing Overhang
In July I wrote Who’s Afraid of Chinese Models, where I argued that nearly everyone’s understanding of the threat Chinese models posed to the frontier labs was overstated, and an artifact of demand exceeding supply. A world with sufficient compute is one where intelligence is a commodity, and in commodity markets margin comes from a superior cost structure, which I would expect the leading model providers to have, in part because they have a meaningful lead in scaling and can apply superior AI to their infrastructure. I wrote:
All of this is to say that I think the reaction to Kimi and Chinese models generally is pretty over-blown, at least from an economic perspective. Right now there is a price umbrella that is downstream of the lack of compute; I highly doubt that Chinese models are cheaper to serve on a marginal cost basis, they just seem cheaper because Anthropic and OpenAI are so supply constrained that they are charging far more than they would if there were sufficient supply to meet the demand for intelligence.
A price umbrella is another way to say price overhang. OpenAI and Anthropic charge high prices because they can: there is so much demand for their products at their current price levels that they can barely keep up. At the same time, a huge amount of OpenAI and Anthropic’s supply does not go towards inference, but rather training, reinforcement learning, and R&D: that is compute that does not go towards reducing their price overhang. If they slowed down progress they could re-allocate computing and fully capture the market.
The Capital Overhang
I reiterated above that I don’t think that there is a bubble in terms of compute availability, but I noted in Nvidia’s Risky Business that there might be a timing problem:
This might not cost Nvidia anything in the end: if AI revenues truly take off, then the debt markets will open back up, and ultimately companies will go back to funding infrastructure investment through free cash flows. Right now, however, is the danger zone, as hyperscalers blow through the debt markets and Google at least starts to tap equity. To the extent Nvidia competes through novel funding mechanisms that, at the end of the day, draw on things like insurance floats and pension funds and other long-run liabilities that are the bread and butter of the asset managers the company is partnering with, the risk — unmarked, unlike equity — is considerably higher.
That’s why I started with 1870 and Cooke’s ill-fated agreement with Northern Pacific. Yes, the upside the deal afforded Cooke was incredible, but it was incredible for a reason: it was very risky, and pioneering new funding mechanisms only served to spread the pain when it all blew up. It’s one thing to spend all of your free cash flow; it’s another thing to tap the debt markets. And, beyond that, it’s a completely new nerve-racking thing to bring safety-seeking assets to bear. AI better deliver before it’s too late.
Anthropic is telling investors that it is profitable, but that comes with a big caveat; from the Financial Times:
Anthropic has told its backers it will be profitable this quarter, as it moves to allay investor concerns about the aggressive cash burn of frontier AI companies ahead of its blockbuster initial public offering. The company has told a small group of shareholders that its adjusted operating income will be positive for the second consecutive quarter, according to multiple people with knowledge of the matter. The measure strips out costs including stock-based compensation. Anthropic’s gross margins are above 80 per cent before accounting for revenue shared with distribution partners, including Amazon, and the cost of training its models, according to two of the people.
Excluding stock-based compensation is a pretty big caveat, but to be fair, the concern in terms of a capital overhang is cash; the bigger issue is the exclusion of training costs, which are massive (and which, in a true depiction of gross margins, would count as depreciation). That is the part that needs to be covered by revenues before the world runs out of capital, which is to say that Anthropic would be fine if only they didn’t have to pay for training! [CORRECTION: I am told that Anthropic is indeed profitable including training costs.]
As it stands, both Anthropic and OpenAI and the hyperscalers and neoclouds need AI revenue to increase dramatically; there is a limit to available capital, and finding that limit will end badly for everyone depending on outside capital to fund future buildouts — even if it by no means would mean the end of AI.
The Safety Overhang
I made pretty clear at the beginning that I fundamentally disagree with the framework that many in the AI safety movement operate under; that doesn’t mean I don’t think that AI safety is a serious issue, or that there aren’t already real risks. Indeed, one of the implications of there being a capability overhang is that cybersecurity risks are very real, and are going to be a massive challenge in the next few years. From Autonomy and Innovation:
The expected value for a hacker’s automated attack is always positive. If the offensive agent finds a vulnerability and creates an exploit, and that exploit fails or is itself buggy, then nothing has changed about the status quo: the exploit doesn’t work (or, perversely, makes the original vulnerability larger by virtue of its own bugs); if the agent executes the exploit perfectly, meanwhile, the attacker has gained access to the system. The attack only needs to work once for the entire endeavor to have a positive payoff.
The challenge for the defender, on the other hand, is that they need to keep the software in question working correctly, and not make the situation worse. This means that any automation has a negative expected value: successful automated vulnerability discovery and patching preserves the status quo, i.e. the software is not hacked. However, any unsuccessful patches make the situation worse, either by breaking the software or by introducing new vulnerabilities. The agent only needs to fail once for the entire endeavor to have a negative payoff.
This is the dynamic that leads to the exact situation Dalton describes, where offensive actors are fully automated while defensive systems, even if they use AI, will be incentivized to keep a human in the loop, and no human in the loop will be able to keep up with fully automated agents. Truly effective defense will mean truly trusting agents to act autonomously, but most companies won’t do that until they are forced to by regular and unremitting hacks by fully autonomous attackers.
First, it’s worth pointing out that this risk is not alignment risk, at least as that term was traditionally defined: a model being directed to do bad things is aligned; suggesting that models ought to know what is good or bad is an entirely different consideration, and I think it is very problematic that these questions have been conflated. The fact of the matter is that LLMs are, if anything, too obsequious; there simply isn’t any evidence of LLMs having a will or operating with malevolence.
Second, arguing against progress because of cyber risk was relevant before the agentic paradigm; at this point the genie is out of the bottle — and open weights models capable of attacks are already here.
Third, this reality actually makes the case for pushing the frontier, not pacing it. The capability overhang is entirely on the offensive side: it’s defenders that need models that are not only good enough to mount a defense, but to do so in an entirely automated way that doesn’t bring down the infrastructure being attacked. We’re not there yet.
In other words, when it comes to the tangible safety risk that exists today — bad actors using aligned LLMs to attack infrastructure — pacing the frontier actually increases the window in which bad things can happen.
This isn’t, of course, what Amodei and the doomers are talking about: they are worried about recursive self improvement enabling AI to improve itself, independent of human control, and I’m open to debating the risks that might result were there actually a tolerance for debate, instead of an insistence on acquiescence.
And, of course, all of the frontier labs are free to pace their own progress. That they won’t is a reminder that this is personal: Anthropic exists because Amodei and his cofounders didn’t trust Sam Altman and OpenAI; I’m not sure it’s a coincidence the demand to pace the frontier came when the frontier was, for the first time in a while, set by OpenAI.
Indeed, that’s probably the overhang that matters most of all: competition. Anthropic is fine with Anthropic being in the lead; anyone else requires government intervention. That there are safety arguments to be made that just so happen to align with their need for more time to build a moat is, I’m sure, but a sign from Silicon Valley’s newest god to its self-ordained priesthood.
Xiaomi open-sources MiMo-V2.6 Pro and Flash models
Xiaomi has open-sourced its omnimodal MiMo-V2.6 Pro and Flash models, which are capable of coordinating agents for complex tasks like 3D scene construction and robotics.
Summary
Deep Dive
- MiMo-V2.6-Pro is the flagship; MiMo-V2.6-Flash is optimized for cost-efficiency.
- Uses reinforcement learning on verifiable complex tasks, completing 750,000 trajectories in six days.
- Pro-UltraSpeed variant offers 20x output speed for latency-sensitive applications.
- Supports Lean 4 formal theorem proving and PFAS chemical research.
- Available on MiMo API, OpenRouter, and MiMo Desktop.
Decoder
- Omnimodal: A model capable of processing and generating multiple types of data (text, image, code, audio, 3D assets) natively without separate specialized sub-models.
- RL (Reinforcement Learning): A training method where models learn by interacting with an environment and receiving feedback (rewards) for their actions.
Original Article
Xiaomi has released and open-sourced the MiMo-V2.6 series, introducing two natively omnimodal models for coding, visual tasks and computer use. MiMo-V2.6-Pro is its most capable model, while MiMo-V2.6-Flash targets a lower-cost balance of intelligence and efficiency. Pro-UltraSpeed offers up to 20 times faster output at the same quality for latency-sensitive work.
MiMo-V2.6-Pro scored 46.32 on the Artificial Analysis Intelligence Index v4.3. Xiaomi says this puts it ahead of Kimi K3 and Qwen3.8 Max as the highest-scoring open-source model in the comparison. API prices stay at V2.5 levels, with Flash priced at $0.14 per million uncached input tokens and $0.28 for output, and Pro at $0.435 and $0.87. UltraSpeed costs ten times more.
The release centers on Xiaomi’s effort to scale reinforcement learning on verifiable, complex tasks. In under six days, Flash and Pro each completed 30 RL steps across roughly 750,000 trajectories, costing about $850,000 and $2.62 million. DeepSWE v1.1 scores rose from 48.8 to 65.68 and from 58.4 to 72.57. Training spanned coding, general-agent, visual and cybersecurity tasks, using 1,568 samples per update and context lengths up to one million tokens. Xiaomi froze the router to limit drift and used adversarial evaluation, anomaly detection and verifier cross-checks against reward hacking.
MiMo-V2.6 moves beyond conventional software work into what Xiaomi calls “Vibe World.” From an image, video or text prompt, it can coordinate agents to construct and visually test interactive 3D scenes. It can create Blender assets, control a Franka Panda robotic arm from camera feeds, produce frontends and presentations, assemble videos, and compose music as scores and MIDI. Research demonstrations included screening materials for capturing PFAS chemicals and helping formalize a Lean 4 theorem in more than 6,000 lines of kernel-verified code.
Pro and Flash are available now in AI Studio, MiMo Code, MiMo Desktop, Xiaomi’s MiMo API Platform and OpenRouter. MiMo Desktop is leaving early access with both models included, while UltraSpeed is offered for real-time workflows. Xiaomi is publishing the technical report, training environments and RL code alongside the models, framing the launch as a reproducible test of scaled RL and model self-improvement.
Introducing Grok 4.7
SpaceXAI launched Grok 4.7, emphasizing improved coding capability, self-verification, and a new safeguard stack designed to prevent risky command execution.
Summary
Deep Dive
- Grok 4.7 outperforms Grok 4.6 in CursorBench 4.0, DeepSWE v1.1, and Terminal-Bench 4.0.
- Features a new 'fast variant' at twice the output speed but double the price.
- Topped LatchBio’s biosafety benchmarks with 62.4% score.
- Restricts risky dual-use prompts while minimizing blocks for legitimate security work.
- Natively trained to handle the Grok Bot harness for conversational tasks.
Decoder
- Jailbreak: Techniques used to bypass safety filters in AI models to force them to perform prohibited actions.
- Dual-use: Technologies that have legitimate civilian applications but can also be misused for malicious purposes (e.g., cybersecurity or biological research).
Original Article
Introducing Grok 4.7
SpaceXAI's most powerful model for coding and knowledge work. Twice as fast, at half the price of comparable models.
Grok 4.7 is our most capable model for coding and knowledge work. It works longer on difficult tasks, checks its own work more carefully, and comes with our best-calibrated safeguards to date. Served at the same price and speed as Grok 4.6, it is highly competitive in its class.
On CursorBench 4.0, which stresses longer-running coding tasks, Grok 4.7 is at the frontier in price-performance.
Model Improvements
Grok 4.7 uses a new, larger base model compared to Grok 4.6. It was trained with a longer reinforcement learning run on a harder mix of tasks, weighted toward problems that take many hours to complete. The model is better at verifying its own work and managing longer context. We also trained Grok 4.7 to natively understand the Grok Bot harness, making it better at conversational tasks and general knowledge work.
Grok 4.7 is better at creating documents and presentations. In GDPval and AA Briefcase, AI is asked to work on tasks done by professionals such as lawyers, nurses, and financial analysts. Grok 4.7 improves upon Grok 4.6 on both benchmarks and performs comparably to other frontier models.
Safety & Cybersecurity
Grok 4.7 was built with an entirely new safeguard stack. It is the strongest model we’ve tested on refusals and jailbreak resistance. In dual-use domains like cybersecurity and biological work, it leads on both utility for benign tasks and safe refusal on dangerous ones, topping LatchBio’s biosafety benchmark at 62.4%.
Grok 4.7 balances strong cyber defense capabilities with low refusal rates for legitimate use. It shows the highest safety on HackerBench v0.3, our benchmark for risky and malicious cyber tasks, allowing only 3.3% of risky dual-use prompts through while rarely blocking legitimate security work. We’ve also started giving select cybersecurity partners invite-only access to Grok 4.7’s red-team capabilities for defense research.
Pricing and availability
Grok 4.7 is available today in Cursor and Grok Build. It is also available through the Grok API, third-party coding harnesses, and model routers and cloud platforms.
The model is priced starting at $2 per million input tokens and $6 per million output tokens. We also serve a fast variant with twice the output speed at twice the price.
Try it in Grok Build for free
Get started today at x.ai/build.
$ curl -fsSL https://x.ai/cli/install.sh | bashThe current balance of power in open models
Chinese labs have established clear leadership in the open-weight AI ecosystem, capturing double the Hugging Face downloads of American models.
Summary
Deep Dive
- Open-weight models are now economically viable, shifting from research projects to production backends.
- Distillation from American models accounts for only 1-2 months of the performance gap, contradicting the idea that Chinese success is solely based on copying.
- OpenRouter data shows open model token usage grew from 1T to 80T per week since September 2025.
- Academic adoption of open models rose from 2% in 2023 to 50% in 2026.
- Chinese labs prioritize specific agentic coding tasks, often outperforming U.S. models on targeted benchmarks.
- Ecosystem reliance on foreign models introduces regulatory and security risks that cannot be solved by simple access restrictions.
Decoder
- Open-weight model: An AI model where the model weights are made public, allowing users to run the model on their own hardware or via third-party inference providers, distinct from fully open-source models which also share training data and code.
- Distillation: A process where a smaller student model is trained to mimic the behavior and outputs of a larger, more capable teacher model to improve efficiency.
- Inference: The process of running a trained AI model to generate predictions or content based on new input.
Original Article
The current balance of power in open models
The expanded form of a testimony I prepared for Congress.
I was recently invited to brief a group of Congressional members and staff on the state of open-weight models in the lens of U.S.-China competition. I’m sharing my prepared remarks as a state of the union on open models that is accessible to a broader audience.
Recap: What is an open source v. open-weight vs. closed model?
Open language models are AI models where their weights are publicly available for inspection or downstream use. These are most often contrasted to so-called “closed” AI models. Closed models offer access only through Application Programming Interfaces (APIs) that developers can use to directly query a model, like GPT-4 or Claude Opus 4.5, or through products, like ChatGPT and Claude Code.
Open language models primarily are bucketed into two categories, open-weight and open-source models. Open-weight models are the most common form, such as popular models like Meta’s Llama, Alibaba’s Qwen, Google’s Gemma, or DeepSeek’s models. These models are governed by licenses, governing documents dictating what is allowed with downstream use, and are often accompanied by inference code in libraries such as Transformers, VLLM, SGLANG, etc. Since about April 2025, Chinese AI companies have been the clear leader in open-weight models.
True “open-source” models are similar to these, as they include the weights, licenses, and inference code, but they also include the complete information needed to reproduce the model – the training code and training data. The most prominent open-source models have been built in the United States, led recently by the Allen Institute for AI’s Olmo models that I helped build in my recent 2.5 years there. The other prominent open-source models are also built by American non-profit organizations, including OpenAthena’s Marin models and EleutherAI’s Pythia models.
Open-weight, open-source, and every other label for a model – including closed models primarily offered via an API – exist on a spectrum. For example, Nvidia’s Nemotron models are far more open than most open-weight models, releasing large quantities of their training data under permissive licenses, but they’re not fully open-source because they do not release all of the data. Closed models also exist on a spectrum based on what information the API reveals and the terms of use.
The state of competition between American and Chinese open-weight models (unit economics, technical capabilities, etc.)
We are living in a world where GLM-5.2 and Kimi K3, some of the latest, leading Chinese models, have enacted a step change in the commercial viability of open models — crossing a similar threshold in agentic capabilities that Anthropic’s Claude Code crossed in December of 2025.
America was the early leader in open language models, primarily through Meta’s Llama models, which were used extensively across research and commercial tasks. Chinese open-weight models surpassed American open-weight models in these two key areas about 18 months ago. The simple metric showing this is Hugging Face Downloads, where China took the lead in July of 2025 primarily through the success of Alibaba’s Qwen models. I personally maintain tools to track this data, and since I first published the American Truly Open Models (ATOM) Project in August of 2025, China’s download lead has grown to about 1.6B – with a total of 3.2B downloads, twice that of America’s total.
On popular capabilities benchmarks, such as the Artificial Analysis Intelligence Index (AAII), the Chinese open-weight models have a clear lead over American counterparts. The top three Chinese models as of writing this on September 14, 2026 are Z.ai’s GLM-5.3 and GLM-5.3-Flash and Moonshot AI’s Kimi K3 with scores of 45, 42, and 44 respectively. By comparison, the leading American models are Thinking Machines’ Inkling and Inkling Small, both with a score of 26, and Nvidia’s Nemotron 3 Ultra, with a score of 23. The top American models were released in June and July of 2026, and are updated less frequently than their Chinese counterparts. For example, Chinese labs released models with scores above these American models 2-6 months before the American companies got there (e.g. GLM-5 or DeepSeek V4 Pro). There is a trend of more American companies releasing models, including names like Arcee AI, Poolside and IBM, but they are not rapidly closing this performance gap. Other benchmarks tell a similar story.
Together, Chinese open-weight models are approximately 2-5 months behind the closed American frontier, with the open-weight American models being approximately 6-9 months behind the likes of OpenAI and Anthropic. The Chinese labs are closest in tasks with clear user demand, such as agentic coding, and further behind on more open-ended scientific tasks, such as physics or biology.
The reasons why Chinese labs can produce these strong models, despite having fewer resources than American counterparts, is still an open debate and heavily influenced by different work cultures, but is also influenced by a few key technical factors. The Chinese labs release their models faster and focus on a slightly narrower distribution of tasks, flattering them slightly on public benchmarks. Releasing faster helps them score higher because all the labs are making consistent progress, so once you “finish” a model to be released, it is a snapshot of performance at that given time — labs where that time is later tend to score higher. Still, the models built by the Chinese labs are genuinely strong and represent real competition to the American industry. This competition will not decrease meaningfully as the closed labs patch vulnerabilities in their API offerings which enable distillation.
Distillation is most impactful in new domains and does not make it trivial to create a universally strong final model. I estimate that if distillation was fully prevented, e.g. with know-your-customer (KYC) tools at Anthropic and OpenAI, the gap from the strongest American models to Chinese open-weight models would only increase by 1-2 months.
For example, the Chinese labs are rapidly changing their posture towards paying for training data in 2026. Earlier in the year, the top Chinese labs including Moonshot AI and Z.ai had a strong preference towards building data workflows in-house, but by the summer they had begun to buy the cutting edge data – challenging RL environments for agentic tasks – from both established American companies and new Chinese startups.
With the advance of open weight models in China towards the frontier of capabilities, and the recent documentation of growing risks around frontier models in areas such as cybersecurity (e.g. the OpenAI-HuggingFace incident), there’s growing regulatory uncertainty on how continued releases can enable a safer ecosystem?
A structural challenge in open-weight models is that there are few effective methods for stopping pieces of open software from reaching bad actors. If an attempt was made to restrict access to the strongest open-weight models from China because they amplify risks, the parties who would be set back are American businesses. We have an example of this – HuggingFace used a Chinese open-weight model to understand the cyberattack because closed models would not answer their requests. Thus, managing the risks of open-weight models often comes down to ecosystem preparation.
Open-weight models are becoming an essential tool for AI diffusion, and the best path to get ahead of these risks and unbalanced relationships where American companies rely on models built in China is to continue to enable investment in open models in the US. Ownership of open models allows better coordination and preparation of risks that are global in their nature while accelerating diffusion of AI services throughout the domestic economy.
The state of open model adoption: How is open-source being used by academia, businesses, and other countries?
Open-weight language models have grown substantially in general interest and economic viability in 2026, allowing early glimpses of more direct ways to compare adoption of models from the US, China, or elsewhere on top of Hugging Face metrics. One example is OpenRouter usage. OpenRouter is a popular LLM inference platform that supplies a single interface to switch between models, open and closed, from the US and China. This platform is primarily known for trying different open-weight models. The platform has shared usage data for the top models since Jan. 1, 2025, and shown growth in usage from ~1T tokens processed from open models in a week of September 2025 to ~80T tokens per week today. In that time, Chinese models have grown from ~70% market share to over 80% of usage. Other platforms that are designed to commercialize open models show similar data, such as the open-source coding agent OpenCode, which shows an inference volume of ~95% or higher with Chinese models.
These open platforms are the best approximation of open model usage we have – a large proportion of open model usage is on platforms that do not disclose per-model breakdowns, such as Together AI or Fireworks AI, and in private deployments for enterprise applications.
Many prominent technology companies and startups have been building on Chinese open-weight models for their AI features, such as Harvey, the legal agent, Cursor, the coding agent, and DoorDash’s use of Kimi models, Airbnb’s use of Qwen, or Perplexity’s use of DeepSeek. These prominent companies are the tip of the iceberg, where a large swath of younger Silicon Valley startups are building on Chinese models in order to have low-cost, flexible options. There is a growing trend of American startups and companies entering enterprise agreements with Chinese model labs in order to get permission to use their models in their products – a new form of cross-border technology collaboration I have not witnessed in my career.
The foundation of innovation on Chinese models extends further into the AI ecosystem. To a first order approximation, most of academic research is conducted on Alibaba’s Qwen family of models. Having met multiple members of the Qwen leadership team during my trip to China, they are very invested in and intentional about this type of adoption, which will not be easy to claw back to American models.
To quantify the adoption of open models across academia, I scanned every paper in the 5 most popular ML categories of arXiv (cs.AI, cs.CL, cs.CV, cs.LG, stat.ML), the preprint platform popular in AI research. The results clearly track my understanding of the evolving leadership in AI research, showing LLMs becoming a foundational layer of ML research – mentions of any open model were 2% in January of 2023 and 50% in September of 2026 – and the leading role shift from the U.S. to China in the same time period.
For example, in April to May of 2023, a few months after Meta’s original Llama (a backronym, Large Language Model Meta AI, first released in Feb. of 2023), about 2,600 of 12,000 new AI/ML papers on arXiv mentioned at least one prominent open model family. Of all those scanned papers, ~5.5% mentioned Llama and ~1% mentioned a Chinese model. In the fall of 2024, during Llama’s peak, about 23% of papers mentioned Llama with about 7.5% mentioning Qwen, the most direct Chinese competition. Today, Llama has lost its lead in academia, being mentioned in about 21% of papers still, which is remarkable longevity, but Qwen’s share has risen to 30% of papers. Overall, any Chinese open weight model is mentioned in over 40% of papers, over the U.S.’s 30%, with China’s share continuing to grow.
This shows that we clearly have a lot of work to do in order to re-establish the U.S. as the home of AI research in the era of open-weight language models. There are signs of hope.
In our research, we find that American models of comparable capabilities-to-size regions to their Chinese counterparts get adopted at disproportionate rates. In the last year we’ve seen OpenAI’s first open-weight models since ChatGPT, gpt-oss, become one of the most adopted open-weight models of all time. Since then, Google’s Gemma 4 models have been some of the only ones ever to show similar adoption numbers to Qwen’s most popular small models, and Nvidia’s Nemotron models have modest adoption despite numerous more capable models at the same size point.
Summary
The story of open models in 2026 is one of establishing economic relevance. This is the convergence of many stories across the AI ecosystem, summarized as:
- The capabilities gap from open to closed models available to users has been decreasing over the last 3 years. This varies by task, but can be estimated as a 2-5 month gap in capabilities. With capabilities overall progressing so fast, this has seen open-weight AI models unlock substantial markets in 2026 and points to more inflection points in the near future.
- Open model usage is exploding in high-value industries (e.g. software engineering, legal services, financial services), indicating an emergence of an alternative ecosystem to the best closed models. Platforms offering inference primarily on open models, from Together, OpenRouter, Fireworks, Baseten, etc., are seeing incredible growth as the first winners of an open model post-training economy (other layers include finetuning APIs such as Thinking Machines’ Tinker). This is combined with numerous anecdotes from technical staff in the AI industry that uses open-weight models such as GLM-5.3 as an alternative to Claude or GPT due to a combination of speed, lower prices, customizable offerings, and privacy.
- Chinese AI companies are the clear leaders in open weight models. Relative to 2025, where Chinese models like DeepSeek R1 shook the AI world with surprise, the American AI labs have been recovering in their positions with open-weight models, but despite more substantial investment in the US, the Chinese labs regularly are producing notably stronger models adored by many types of users.
- Distillation of American AI models by Chinese labs does not explain the entire story of their success. Distillation is an industry standard technique of training another AI model on the outputs from a usually stronger model. The technique is most prevalent in the Chinese AI industry, which has used basic exploits to extract reasoning traces and additional data from American companies’ products that are not fully secured. The best estimates are that distillation helps reduce the performance gap of Chinese companies relative to the American frontier by 1-2 months.
- Chinese models, particularly Alibaba’s Qwen family, are established as a foundational layer of research and development across academia and local model users. In recent months, Chinese open weight models were mentioned in 38% of AI papers, above the U.S.’s 28% – and the Chinese share is growing much faster than its American counterparts. This, along with other political factors and the closed nature of leading American AI companies, is contributing to an accelerated decline in America’s lead as the preeminent AI research hub in the world.
- Open weight models are entering the capability levels where new risks, e.g. cybersecurity, can be enabled by numerous open-weight models being available, necessitating an ecosystem level response in preparation. This new era of risks is also enabling a period of political uncertainty, where there is regulatory attention on the strongest AI models, but massive uncertainty on how policy would be legally enacted. At the same time, many researchers and engineers rely on open models due to more permissive safeguards, where the closed models such as Claude and GPT often refuse critical cybersecurity defensive work or biology research.
Conclusions
In 2026 the Chinese labs are clearly maintaining their status as the leaders of the open-weight AI ecosystem. This comes as open-weight models have passed an inflection point in economic viability and in the face of increased activity from American labs as model competition. The leading Chinese labs do not appear to be meaningfully challenged, as they expand their enterprise and research adoption globally.
This landscape of open models comes at a crucial time in the broader AI ecosystem. We’re seeing OpenAI and Anthropic take massive steps forward with their latest public models, and at the same time call for coordinated care on how we manage the next stage of AI progress. What is happening in the confines of a few AI labs today, especially with extreme talent and compute density, is a precursor to what will soon emerge in the open model ecosystem. Open models are going to be the substrate for everyone else in the world outside of the few true frontier AI labs, to harness an acceleration in software engineering and other computational practices. This represents a substantial source of soft power, influence, and potential for the organizations that enable this broad access to transformative intelligence.
With this future coming soon, we need to collectively stay humble about the exact path open models will take. There are a lot of unknowns with open models – e.g. we don’t have good data on how they’re used in countries other than the U.S. and China. With the distribution of ML training expertise being broad, i.e. tens of organizations and thousands of people that are within a year of the frontier of capabilities, it is a matter of when, not if, open models cross the performance thresholds that enable new workflows. The collective approach should be to understand how to use this broadly accessible, open intelligence for good while proactively mitigating the potential harms.
Thank you to Florian Brand and Kevin Xu for feedback and/or suggestions for this work. For more research informing this post, see the open-source AI reading list.
MiMo-V2.6: Scaling Reinforcement Learning Towards Self-Improvement
MiMo-V2.6 introduces Groupwise Advantage Redistribution to prevent reward hacking and keep reinforcement learning stable at scale.
Summary
Deep Dive
- Uses Groupwise Advantage Redistribution (GAR) to rank multiple attempts, penalizing 'hacks' and rewarding high-quality code.
- Training cost for the Pro model reached $2.6M in post-training alone.
- Found that freezing the MoE router during RL is necessary to prevent expert load imbalance.
- Achieved a 72.6% average@3 score on the DeepSWE v1.1 engineering benchmark.
- Supports long-horizon tasks (up to 1M context) across code, visual, and cybersecurity domains.
Decoder
- Mixture-of-Experts (MoE): A model architecture where only a subset of internal 'expert' subnetworks are activated for any given input, improving computational efficiency.
- Reward Hacking: When an AI model finds a way to get a high score on a test without actually performing the task correctly (e.g., fetching a pre-existing answer rather than solving a problem).
Original Article
Abstract
Reinforcement learning (RL) is the central training paradigm for advancing large foundation models towards self-improvement. This report introduces the MiMo-V2.6 series, an omni-modal family that pushes the frontier of model intelligence by scaling RL compute. Prior to RL, we conduct mid-training on a broad multimodal corpus to provide ample exploration space, and build a solid infrastructure on the pretrained hybrid-SWA architecture to support subsequent scale-up. We scale RL compute along three dimensions: (1) larger batches and higher throughput, with an asynchronous training that consumes 1,568 samples and 2.7∼3.7B tokens per step at context lengths of up to 1M; (2) more diverse and complex environments, spanning code, general, visual, and cyber domains under a mixture of agent harnesses; and (3) more grader compute, via groupwise agentic grading that yields more accurate reward signals for long-horizon tasks and steers the model towards shorter, more token-efficient solutions. To keep training stable at scale, we freeze the MoE router and establish a multi-layer defense against reward hacking. We further build infrastructure for mixed-task agentic RL, including a unified trajectory representation, high-concurrency multi-framework rollout, decoupled control and data planes, and training inference consistency. We open-source the training dynamics, RL environments, and RL framework to facilitate reproduction and further research on scaled RL and model self-improvement.
AI Overview
Suppose you are building a tool that lets an AI agent fix real software bugs autonomously. The agent reads a bug report, explores the repository, writes a patch, and runs tests—sometimes across dozens of turns. When a patch passes all tests, you record a reward and move on. But two patches can both pass: one makes the minimal, correct change; the other swallows exceptions, adds broad compatibility branches, or silently fetches the published fix from the internet. A binary test score cannot tell them apart. The authors of MiMo-V2.6 argue that this is not a quirk of one task type but a fundamental limit on how useful reinforcement learning can become for agents. Their answer is to compare sibling attempts within each training batch, shift the learning signal toward the cleaner passing solution, and run this graded feedback across coding, professional workflows, visual design, and cybersecurity tasks simultaneously. Across all four domains, both model sizes improved steadily as reinforcement learning training progressed.
Scaling the practice field
Reinforcement learning for agents works by having the model attempt a task, receiving a signal about how well it did, and adjusting the numbers that control its behavior—its weights—in the direction that produced better outcomes. The harder the task and the longer the interaction, the more attempts and the more accurate feedback the training process needs.
MiMo-V2.6 processes 1,568 prompts per training step, generates 16 attempts per prompt, and accumulates 2.7–3.7B training tokens each step—roughly the text equivalent of several thousand academic papers flowing through a single update. Because some tasks finish in minutes while others run for hours, the infrastructure must prevent fast tasks from flooding the training batch while slow ones are still running. The authors handle this with adaptive scheduling that tracks each task type's typical duration and acceptance rate, reserving enough in-flight slots to keep every domain consistently represented. Trajectories—the full record of what the model said and what tools returned—are stored in a distributed cache and fetched only when needed for training, so no single machine has to hold the entire batch in memory.
The RL post-training expenditures were $2.6M for the larger model and $0.9M for the smaller one, separate from pre-training and other development costs.
A better signal than "pass"
The central mechanism is called Groupwise Advantage Redistribution, or GAR. Here is what it does in concrete terms.
After each training step, the model generates sixteen attempts at the same task. Some pass, some fail. The binary reward treats every passing patch identically. GAR instead places all sixteen attempts in a shared workspace—same problem statement, same repository, all submitted patches visible together—and an evaluator model examines them jointly. It assesses each passing patch on five dimensions: whether the solution approach fits the problem, whether the implementation is precise without unnecessary fallbacks, whether the changes are minimal relative to what was requested, whether anything outside the task scope was accidentally altered, and whether the code style matches the surrounding codebase. Passing patches are ranked; the bottom of the ranking receives less positive reinforcement and the top receives more. A confirmed hack—a patch that earned its score by fetching the published fix from the internet rather than deriving it—is reset to zero and treated as a failure before any rankings are computed.
The offline complement, Groupwise Reward Synthesis, multiplies the binary test reward by quality and behavior scores from precomputed task-specific rubrics, so failed attempts retain zero reward while passing attempts are further separated by implementation quality.
The key question is whether this grading change actually alters what the model learns, rather than just adding overhead. The authors test this directly: the same model, the same code-only RL setup, the same benchmark, with and without online groupwise grading. Without it, the model's average number of turns and total token length both rise rapidly, more and more attempts hit the context length limit before finishing, and improvements in pass rate plateau. With groupwise grading, pass-rate gains continue through step 52 while turn counts remain roughly stable and token length grows more gradually.
The grader itself consumes compute—about 12.7% of MiMo-V2.6-Pro's total RL cost—so this is a trade-off, not a free improvement.
Stability prerequisite: freezing the router
MiMo-V2.6 uses a Mixture-of-Experts architecture, meaning each transformer layer contains many specialized subnetworks and a router that decides which ones each input token should use. During large-scale RL, the authors found that allowing the router's weights to be updated caused a severe imbalance: a few experts absorbed nearly all the traffic while most sat idle. Coefficient of variation of expert load rose from 0.78 to 2.0 over twenty training steps, peak load from 6× to 16× the average, and the fraction of cold experts from 0.5% to 22%. When the team restored the router's original weights at step 20 while keeping everything else unchanged, load balance recovered immediately and benchmark performance was unaffected—showing the collapse was caused by router drift, not by the expert weights themselves.
The solution, freezing the router throughout RL training, is a stability prerequisite. Without it, scaling the batch size would likely make the imbalance worse.
What the run establishes
On DeepSWE v1.1, a benchmark of 113 long-horizon repository-engineering tasks evaluated by functional verifiers across five programming languages, MiMo-V2.6-Pro's average@3 score—the average verifier pass rate across three sampled attempts—rose from 58.4 to 72.6, and Flash from 48.7 to 65.7 over the course of RL. Average@3 counts how often an attempt succeeds on average, not whether any one of three succeeds. The cross-domain pattern holds in professional workflow automation, visual code generation, and cybersecurity vulnerability reproduction, though three of the headline benchmarks are internal evaluations.
To make the approach reproducible, the authors released MiMo-V2.6-Distill-Qwen-9B, a smaller model fine-tuned from Qwen3.5-9B on MiMo-generated demonstrations. Starting from that checkpoint and running RL with the released environments and verifiers, RL improved every reported evaluation across all four domains—for instance, SWE-bench Verified rose from 61.1 to 66.2 and cybersecurity from 31.3 to 47.0 on an internal mini-benchmark. This is evidence that the training package can be reproduced at smaller scale, not that the trillion-parameter results can be replicated cheaply.
The final model comparisons with other frontier systems involve mixed public and internal benchmarks and do not isolate RL's contribution from architecture or pre-training. The main run's improvements also generally accompany rising token use, so the result is sustained capability improvement, not necessarily shorter or cheaper solutions overall.
Feedback is the scale
The opening question was whether binary "pass or fail" feedback can sustain agent improvement as tasks grow longer and more varied. The controlled GAR ablation supports the paper's proposed mechanism: without quality-aware grading, the model learns to write longer, more defensive patches that pass tests while accumulating bad habits; with grading, improvements continue longer and the model stays closer to the requested scope. The broad training curves show that this graded feedback generalizes across heterogeneous tasks within one training run. Together, the evidence reframes what RL scaling means for agents: not just generating more rollouts, but building a larger, more varied, and better supervised learning loop.
Aikido Altar: open-weight AI for sovereign security
Aikido Altar is a compressed, open-weight security model that lets teams run advanced pentesting within their own air-gapped infrastructure.
Summary
Deep Dive
- Uses Cerebras REAP (Router-weighted Expert Activation Pruning) to select domain-specific experts.
- Reduces size from 1.51 TB (BF16) to 328 GB (W4A16).
- Maintains 92% of original vulnerability coverage despite significant pruning.
- Specifically optimized for security-focused workflows like code analysis and vulnerability remediation.
- Requires a 4-H200 GPU node for production-scale serving.
Decoder
- Quantization: The process of reducing the precision of model weights (e.g., from 16-bit to 4-bit) to reduce memory usage.
- Expert Pruning: Removing specialized sub-networks (experts) from a mixture-of-experts model to decrease size and increase speed.
- Sovereign Security: A security strategy where all diagnostic and offensive tools operate locally within the organization's perimeter.
Original Article
Introducing Aikido Altar: the model that makes sovereign security intelligence possible
Today, we’re introducing Altar, our first open-weight security model, built to bring frontier-grade defensive security into infrastructure you control.
We took our first step toward sovereign security intelligence with our autonomous pentesting appliance Aikido Machine, which runs entirely inside a customer’s own infrastructure, including fully air-gapped environments. Teams with strict rules about keeping code in-house can continually detect, exploit, and validate vulnerabilities across their attack surface without breaking those rules.
Altar provides Aikido Machine with advanced AI that defends, without sending an organization’s most sensitive context to a third-party inference service.
To do this, we needed to bridge the deployment gap. Most capable open-weight models remain difficult to deploy at production scale while retaining the reasoning quality of frontier-grade models.
We started with GLM-5.3, one of the strongest models in our security evaluations. The full model occupies 1.51 TB; quantization brings that down to 488 GB, and our expert-pruning process reduces it further to 328 GB, all while preserving most of the reasoning quality of the parent model.
The gap between frontier and open-weights models explained
Closed frontier models run on somebody else’s infrastructure
Using one means your source code, your internal architecture documentation, and your unremediated findings leave your network. For a bank operating under a data-residency mandate, a hospital group subject to strict data-processing rules, or an industrial operator whose OT environment has no route to the internet at all, it is not a trade-off they're allowed to make.
The deployment gap
Serving a model such as GLM 5.3 at full precision requires hundreds of gigabytes of unified memory and yields limited inference speed when serving parallel conversations with high context windows. In human terms, they’re enormous.
Those resource demands come partly from how these models are built. Many of today's strongest models use mixture-of-experts architectures. These models are made up of a multitude of small specialized neural networks, each called an expert. Only a small number of experts are active for each token, while the full expert pool still has to be stored and served. That means paying a significant memory and infrastructure cost for capacity that may contribute little to the workload you're actually running.
Security work, such as reviewing code for vulnerabilities, proposing patches, and running a penetration test, only calls on a small slice of that expert pool. But the memory cost doesn’t shrink to match: every request still has to load the full model, whether or not it is relevant to the job.
Now add agents into the mix, and the usage of context explodes, which saturates memory and leads to agents fighting each other for room. Each one keeps a running record of everything it’s seen and done, and that record sits in GPU memory alongside the model itself, growing the longer an investigation runs and multiplying across every investigation running in parallel.
That’s why model size matters here: the model and all that growing context draw from the same fixed pool of memory, so more of one leaves less room for the other. The goal is to make a frontier model more efficient to operate for a specific workload, preserving the capabilities that matter while freeing capacity for everything running alongside it.
What can we remove without losing what matters?
The heavy burden of carrying so many experts into a model can be turned into an advantage with the right techniques. In order to minimize the size of open-weight models in an agentic security context while maintaining high accuracy, we used quantization and pruning.
Expert pruning removes some of the experts from the model, typically by deleting entire expert weight blocks and adjusting the model so each token only chooses among the ones that remain. The removal itself is mechanical. Deciding which experts to remove is the hard part, and the loss can be uneven: a smaller model may retain strong coding performance while losing much of its ability to understand a particular natural language. That can leave it unable to reliably interpret documentation, business rules, or application features described in that language.
To decide what to keep, we started with the most natural piece of data: traces from our pentesting harness running internal benchmarks. These capture the code, tool calls, and responses that agents work through during a full pentest, giving us representative inputs to guide expert selection. No customer data was ever involved.
This calibration step gives us a workload-specific basis for expert selection, without training new skills into the model.
The harness traces covered the technical workload. We also needed to preserve language understanding: investigating an application means understanding its features, business rules, and intended workflows, including when its documentation or interface is in French, Dutch, or another language. We therefore added multilingual text to help preserve those capabilities during pruning.
Selection matters as much as size
Many pruning techniques exist to reduce the footprint of LLMs, so we decided to go with a technique that can preserve performance for our use cases: Cerebras REAP, Router-weighted Expert Activation Pruning. We used REAP’s contribution score with a domain-preserving aggregation strategy selected through fidelity experiments. Counting how often an expert is selected gives only part of the picture. REAP estimates its contribution using both the router’s weighting and the magnitude of the expert’s output.
We also examine contributions across different groups of examples. Otherwise, a capability that matters to a less common workload can be lost in an overall average. Cybersecurity, coding, and language capabilities are distributed across experts; there is no cleanly labeled set of “cyber experts” to keep or discard.
In the supporting compression study, models retaining the same number of experts differed substantially in how closely their outputs tracked the reference model. Size alone did not determine what survived.
From 1.51 TB to 328 GB
We achieved a total 78.2% model size reduction compared with GLM 5.3 at full precision, and a 32.8% reduction compared with GLM 5.3 with an AWQ INT4 quantization. Altar retains 168 of the original 256 routed experts in each backbone expert layer, removing 88, or 34.4% of them. The router still selects eight per token, now from the smaller bank.
The other part of the reduction comes from quantization. A model’s weights are the numerical values it learns during training. Quantization stores those numbers with fewer bits, trading some precision for smaller stored weights. Unlike pruning, it does not remove experts.
Our starting checkpoint used AWQ to store most expert weights in four bits instead of the BF16 model’s sixteen. AWQ uses information about the model’s activity on example inputs to limit the errors introduced by lower precision. The intermediate values produced during inference, called activations, remain 16-bit: hence W4A16, or four-bit weights and 16-bit activations. Some weights also retain higher precision. We then applied pruning to this already-quantized checkpoint.
Comparing both parent representations shows what each step contributes:
| Checkpoint | Stored weights |
|---|---|
| GLM-5.3, unpruned BF16 (16-bit) | 1,506.7 GB |
| GLM-5.3, unpruned AWQ INT4 | 488.2 GB |
| Altar, pruned W4A16 | 328.0 GB |
That is 78.2% less storage than the full 16-bit model. Against the already-quantized parent, pruning removes another 160 GB, a 32.8% reduction.
The impact of compression on vulnerability identification capabilities
We have previously built an internal CVE benchmark which we use to assess a model's ability to identify complex real-world vulnerabilities using our AI Code Analysis harness: 32 known vulnerabilities across 30 repositories, with three runs per case.
We ran Altar against it. Altar averaged 60.4% recall per run and rediscovered 23 of the 32 vulnerabilities at least once across three runs.
In comparison, the quantized GLM-5.3 AWQ averaged 61.5% recall and covered the same 23 vulnerabilities. The original GLM-5.3 model at full precision averaged 65.6% recall and covered 25 of 32.
In other words, reducing the already-quantized checkpoint from 488 GB to 328 GB, a 32.8% reduction in stored weights, resulted in approximately a one percentage point less average recall compared with the AWQ baseline, while retaining all of its vulnerability coverage. Compared with the original parent, Altar retained 23 of its 25 covered vulnerabilities, or 92%, with a 5.2 percentage-point decrease in average recall. That’s 92% of the parent’s vulnerability coverage kept at 33% less storage.
That's the tradeoff we were looking for: a materially smaller model while preserving most of the parent's security capability.
We report average recall per run separately from coverage across three runs: finding a vulnerability once is different from finding it consistently. Completed runs without a finding count as misses; incomplete runs are reported separately.
This measures targeted CVE rediscovery within a pipeline that uses other models for surrounding stages. It does not measure blind discovery across an entire codebase, execute exploits to validate findings, or evaluate the fix-proposal stage. Those boundaries keep the benchmark distinct from the broader pentesting workflow we are crafting Altar to support.
Additionally, we deployed Altar to our Aikido Machine fleet immediately after completing these evaluations. Shortly after deployment, it identified a valid critical-severity vulnerability during a client production pentest.
What’s next
The underlying idea is broader: sovereign security intelligence should run inside any environment it protects.
Altar is only the beginning. On compression, we’re exploring lower-bit formats such as EXL3, which could let us retain more experts, alongside further H200 serving optimizations.
The next step is to go beyond compression into training: fine-tuning models for security workflows, improving tool use and long-horizon reasoning, and building a self-learning pipeline informed by our internal benchmarks to guide and shape future models.
That work will extend across vulnerability research, code analysis, remediation and other defensive security workflows.
Aikido Labs exists to keep pushing that further: models that get more capable over time without losing the constraint that matters, defenders have to be able to run them entirely inside the environments they're responsible for protecting.
Altar’s weights are available under Aikido’s organization. Altar can be deployed and served comfortably using a 4-H200s node and the latest version of vLLM.
Projects Redesigned: From Folder to Conversation
Claude Code now features a coordinator model that manages parallel, multi-thread tasks to automate complex development workflows across repositories.
Summary
Deep Dive
- The project system now uses a coordinator to scope requests and delegate tasks to parallel threads.
- Each thread operates in its own cloud session on a distinct git branch.
- A shared memory layer tracks context and project history, reducing redundant prompt engineering.
- Claude can now autonomously open pull requests and run tests, managing conflicts like a human developer.
- Users can monitor progress in a main project chat while diving into specific threads for fine-grained control.
- The system is initially available to specific subscribers, with wider enterprise rollouts expected soon.
Decoder
- Agentic Workflow: A mode of operation where an AI model is given a goal and autonomy to execute a series of steps—such as file manipulation, testing, and Git commits—without needing constant human intervention for each action.
- PR (Pull Request): A standard method in software development for submitting changes to a code repository for review before they are merged into the main codebase.
Original Article
Projects redesigned: from folder to conversation
A new experience for Claude projects, now available in beta in Claude Code
Managing multiple sessions across a build used to require you to divide the work, juggle handoffs, and stitch the results back together. Now in a Claude Code project, you describe what needs to get done and Claude manages the work.
Claude scopes the request, delegates the work, coordinates parallel threads, reviews the outputs, and assembles the finished result. You can steer progress throughout, even from your phone, and it keeps working after you step away from your computer.
For example, configure a project and set a goal to reduce your app's checkout p75 latency. Then ask Claude to profile each endpoint, test optimizations, and open PRs in parallel threads. Or connect your API, web, and mobile repos and set a goal to retire a deprecated v1 endpoint. Claude creates a thread per repo to migrate the callers, run the tests, open PRs, and then tells you which ones need to merge first.
Starting today, updated projects are available in beta to select Claude Pro and Max subscribers who use cloud sessions in Claude Code and don’t have any existing projects on the web or desktop.
Over the coming week, we'll expand access to more Claude Code users on those plans. Updated projects across all of Claude and Team and Enterprise plans come after that. If you're on Pro or Max and don't have access yet, you can join the waitlist.
Existing projects on Pro and Max plans keep working as they do today. We'll upgrade them as the rollout expands to chat and Cowork.
Threads do the work, Claude directs it
Projects have threads that do the work and a coordinator that directs them.
When you start a project, you select a goal as well as the repo or context. Claude starts by suggesting work it can pick up right away. You can configure the project’s cloud environment, connectors, plugins, instructions, and model.
You can monitor and guide progress in the main project chat, or dive into each individual thread to examine and steer the details. Brief Claude in the project the way you'd brief a chief of staff and it routes work to new or pre-existing threads.
Claude also checks in and follows through on work. With repositories connected, a thread opens pull requests and runs your tests; with documents, it reads them and drafts.
Under the hood, each thread is a Claude Code cloud session working on its own branch and copy of the repo. The coordinator keeps work organized, but if any threads work on the same code, the overlap is resolved as a merge conflict just like any other PR.
Each thread can further split its delegated work into pieces using subagents, loops, and workflows when needed so large assignments finish faster.
Context builds over time
Projects are designed for long-running or agentic workflows: work that takes longer than one reply and has more than one part.
Over time, Claude learns more about the project details and applies them to its work. Every thread now adds to and draws from a shared memory, reducing the need for complex prompt engineering.
For example, Claude can remember the release moved to Friday, why the export was dropped, or who to check in with before touching the billing service.
Claude also remembers your working and communication style. You can ask it to adjust how often it checks in, how frequently it starts new threads, or how detailed to make each update.
Alongside memory, projects now include a library that collects the files you add and the artifacts produced by Claude. This makes it easier to find relevant materials and for new work to build on past efforts.
What's next
Projects can run several threads at once, and each one is a full Claude Code session. Because of this, projects can reach usage limits faster. You can check project specific usage and select the model and effort levels used by the coordinator chat as well as the worker threads.
Threads run in the cloud today; running on your machine alongside your local tools and code and behind your network is coming very soon.
Start using projects.
AI Can Build Your UI. Now Figma Can Tell It When It Screwed Up
Applitools now allows Figma to serve as the ground-truth visual baseline for AI-driven tests, bridging the gap between design and implementation.
Summary
Decoder
- Visual Baseline: The approved reference image or design frame against which new code or UI iterations are compared to detect visual regressions.
- MCP Server: Model Context Protocol server, which allows AI models to connect to external systems like design tools or testing suites.
Original Article
Applitools' integration lets a Figma frame URL serve as the visual baseline, matching the test viewport and flagging where an implementation drifts from the approved design. The same September 15 release expands its Eyes MCP Server so coding agents can inspect, review, and resolve visual tests from chat. Code that runs is not the same as code that matches the design, and Figma becomes the reference that settles it.
SpaceXAI releases Grok 4.7 for coding and knowledge work
SpaceXAI's new Grok 4.7 model prioritizes longer task execution and self-checking, aiming to compete with frontier models in coding and knowledge work.
Summary
Deep Dive
- Grok 4.7 increases its CursorBench 4.0 score to 46.3% and DeepSWE v1.1 score to 71.0%.
- The model utilizes a new reinforcement learning cycle and a harder task mix.
- It features improved jailbreak resistance and scores 62.4% on LatchBio’s biosafety benchmark.
- The release targets high-effort professional tasks like document analysis and complex coding.
Decoder
- Red-team: The process of intentionally attempting to break or exploit an AI model to identify safety vulnerabilities and weaknesses.
Original Article
SpaceXAI has released Grok 4.7, its most capable model for coding and knowledge work, with access now open in Cursor, Grok Build, the Grok API, third-party coding harnesses, model routers, and cloud platforms. The company positions it as a faster, lower-cost rival to other frontier systems. Pricing starts at $2 per million input tokens and $6 per million output tokens, matching Grok 4.6, while a fast variant offers twice the output speed at twice the price. Grok Build also offers free access to try the model.
Grok 4.7 works longer on difficult tasks, checks its work more carefully, and comes with our strongest safeguards to date.
Grok 4.7 runs on a new, larger base model trained through a longer reinforcement learning cycle and a harder task mix weighted toward work that can take many hours. It is designed to stay on difficult assignments for longer, check its output more carefully, and manage extended context. Native training on the Grok Bot harness also targets conversational tasks and general knowledge work.
The benchmark results show clear gains over Grok 4.6. Grok 4.7 scored 46.3% on CursorBench 4.0, up from 40.4%, and reached 71.0% on DeepSWE v1.1 at high effort. It also posted 64.0% on EEBench, 1,657 on AA Briefcase v1.1, 38.0% on Terminal-Bench 4.0, and 19.6% on the Harvey Legal Agent Benchmark. Its 56.7% HealthBench Professional score trailed GPT-5.6 Sol Max and Fable 5.1 Max, showing that its lead is not universal. Results from GDPval and AA Briefcase point to stronger document and presentation work across professional tasks for lawyers, nurses, and financial analysts.
Safety is another major part of the release. SpaceXAI says Grok 4.7 uses an entirely new safeguard stack and is its strongest model yet for refusals and jailbreak resistance. It scored 62.4% on LatchBio’s biosafety benchmark and allowed only 3.3% of risky dual-use prompts through on HackerBench v0.3, while rarely rejecting legitimate security work. Select cybersecurity partners are also receiving invite-only access to its red-team capabilities for defense research.
For SpaceXAI, Grok 4.7 ties the Grok model line more closely to its own coding products while extending access across external developer tools and infrastructure. The combination of longer task execution, self-checking, competitive token pricing, and tighter safeguards makes the release a direct push for coding teams and knowledge workers choosing among frontier models.
Googlebooks launch October 4 starting at $899—here are the five models you can preorder today
Google is attempting to break into the premium laptop market with 'Googlebooks,' Android-powered machines featuring deep Gemini AI integration and high-end hardware.
Summary
Deep Dive
- Models start at $899, with most configurations priced above $1,000.
- Gemini is integrated through 'Magic Pointer' and 'Rambler' voice tools.
- The OS allows native Android app execution without ChromeOS virtualization.
- Sideloading requires developer identity verification.
- All models come with one year of Google AI Pro (5TB storage) and GeForce Now.
Decoder
- OEM (Original Equipment Manufacturer): Companies that manufacture hardware components or products that are then rebranded and sold by other companies.
- TOPS (Trillion Operations Per Second): A metric used to measure the performance of an AI accelerator or NPU.
Original Article
Google is ready to take a swing at laptops… again. The company has had a lot of success in phones, but it has struggled to parlay that into an effective strategy for larger screens. Chromebooks have been fine as a budget option, but it seems like Google has accepted they’ll never compete with macOS or Windows. That’s where Googlebooks come in. These new Android-powered laptops are intended to be “real” computers, featuring premium hardware, on-device apps, and powerful development tools.
After the Googlebook reveal at I/O this year, Google’s partners are now ready to sell you one. The first Googlebooks are available from several OEMs that have been making Chromebooks over the years, but don’t expect Chromebook pricing. There’s only one model under a grand, and the rest are well above that.
Familiar hardware
The bulk of Chromebooks have been budget machines, and they looked the part. Googlebooks, however, feature bright OLED screens, metal builds, big haptic trackpads, and the Glowbar. All the machines have this illuminated bar on the lid, but it doesn’t do much for now. It lights up when the device boots and can display battery level, but Google says it’s going to release an API for developers to add more features.
Across the full lineup, Google says you can expect 16GB of RAM with 32GB upgrades for some models. At launch, the machines will run either a Snapdragon X Elite or an Intel Core Ultra Series 3 (Panther Lake) chip. Both platforms have an NPU capable of 45 TOPS of on-device AI performance, and all the designs are fanless. Regardless of the exact model, Google says you can expect up to 14 hours of video playback and 16 hours of web browsing.
Here are all the launch models and their starting prices.
| Acer Googlebook 14 (starting at $899) | |
|---|---|
| CPU | Intel Core Ultra 5/7 (Panther Lake) |
| Memory | 16 GB | 256 GB+ |
| Display | 14-inch 2.8k OLED touch display, 500 nits |
| Measurements | 13.9mm, 2.7 lbs (1.22 kg) |
| Design | Aluminum and Magnesium chassis |
| Biometrics | Fingerprint sensor |
| Dell XPS Googlebook (starting at $999) | |
|---|---|
| CPU | Qualcomm Snapdragon X Elite |
| Memory | 16-32 GB | 512 GB+ |
| Display | 13.4-inch 2.5k IPS touch display, 500 nits |
| Measurements | 12.5mm, 2.5 lbs (1.13 kg) |
| Design | Machined Aluminum chassis, Haptic touchpad |
| Biometrics | Fingerprint sensor |
| HP Googlebook 14 (starting at $1,299) | |
|---|---|
| CPU | Qualcomm Snapdragon X Elite |
| Memory | 16-32 GB | 256 GB+ |
| Display | 14-inch 2.8k OLED touch display, 500 nits |
| Measurements | 14mm, 2.7 lbs (1.22 kg) |
| Design | All-Aluminum build, Haptic touchpad |
| Biometrics | Fingerprint sensor |
| Lenovo Googlebook 15 (starting at $1,299) | |
|---|---|
| CPU | Intel Core Ultra 5 (Panther Lake) |
| Memory | 16 GB | 256 GB+ |
| Display | 15.3-inch 2.8k OLED touch display |
| Measurements | 15.95mm, 2.68 lbs (1.22 kg) |
| Design | Magnesium-alloy and carbon-fiber chassis, Haptic touchpad |
| Biometrics | Face auth and fingerprint sensor |
| Asus Googlebook 14 (starting at $1,299) | |
|---|---|
| CPU | Intel Core Ultra 5/7 (Panther Lake) |
| Memory | 16-32 GB | 256 GB+ |
| Display | 14-inch 3k OLED touch display, 500 nits |
| Measurements | 14.0mm, 2.2 lbs (1.0 kg) |
| Design | Aluminum and Magnesium chassis, Haptic touchpad |
| Biometrics | Fingerprint sensor |
The Acer model is the cheapest, and it’s also the only convertible in the lineup. The price jump from that device to the next cheapest (the Dell XPS) is a substantial $300 if purchased from Google, and the rest of the launch laptops are $100 more than that. Interestingly, the Dell XPS Googlebook is $200 cheaper ($999) if ordered through Best Buy. We’ve asked Google about why the pricing discrepancy is so wide.
Google’s partners launched a few Chromebooks in this price range, but they didn’t see wide adoption.
The apps you know but bigger
A major part of Chrome’s failure to truly compete with macOS and Windows is down to software availability. Google tried multiple ways to shoehorn local apps into its web-first platform, but it never took. The advantage of starting from scratch, according to Google, is that these new laptops share the same software stack as Android phones. That means you can run Android apps natively without all the VM weirdness of ChromeOS, and there are a ton of those.
Google is aware of the sordid history of running phone apps on desktop PCs. It’s often a clunky experience using a phone app designed for touch on a large screen with a keyboard and mouse, but the company claims you may be surprised how many apps in the Play Store are now optimized for laptops. The Play Store client will call out apps and games that have been optimized for bigger displays, and Google says more apps are being updated for laptops all the time. There will also be a “desktop-class” version of the Chrome browser preinstalled on Googlebooks.
Even if optimized, these are still going to be apps designed for, and primarily used on, smartphones. It’s unclear if they’ll offer the kind of power people expect for a $1,300 laptop. On the gaming side, Google notes that its Level Up program is delivering more experiences suitable for a keyboard and mouse, but the selection of desktop-ready games won’t be anything like what you get on Windows or macOS.
There is, of course, nothing stopping you from using phone apps on a Googlebook if you want. They’ll run in a window, and you may not even need to install them. With Cast My Apps, the apps installed on your Android phone are available instantly on a Googlebook. They appear in a dedicated app list on the laptop, allowing you to open and use them without touching the phone. The file manager can also seamlessly access files on your Android phone.
Similarly, a feature called Continue On lets you pick up where you were in a phone app on the Googlebook version of the app. This requires developer support, so not all apps will carry over. Still, this level of integration with Android is not something you get on either Windows or macOS, which could make a Googlebook appealing if you’ve got a Pixel or Galaxy riding around in your pocket.
As you would expect, the Play Store is the primary source of apps for Googlebooks, and there will be some limits on sideloading. Google has confirmed to Ars that Googlebooks will enforce developer verification requirements for apps similar to Android phones. Under this system, only apps from developers who have verified their identities with Google will be eligible for installation from a downloaded APK file. This requirement will roll out broadly next year, but it’s being activated in a few markets in the coming weeks.
The AI play
Google is following through on its promise to make AI a core part of the Googlebook experience. Gemini models are going to be deeply integrated with the software, ready to work on your screen context with a cursor wiggle. Google calls this Magic Pointer, and it has a few suggestions about how to use it.
Let’s say you have an email or PDF with dates you need to add to your calendar. You can jiggle the cursor, point Gemini at the email, and get all the dates added in one go. If you get a suspicious email, you can invoke Gemini with your cursor and ask it to figure out if it’s a scam. You can also point Gemini at images and tell it to edit them or generate new versions without downloading anything. Useful? Maybe occasionally, but Google’s enthusiasm for Magic Pointer seems a bit premature.
Rambler, the AI-powered voice input tool that debuted on Pixel 11 phones a few weeks ago, will also be included on Googlebooks. That means you can start talking to fill text in any field, with streamlined instant editing to remove “ums” and corrections. On phones, I’ve found this can make your writing voice sound like a sterile AI output, but if you’re just banging out a quick email, maybe that’s fine for most people.
AI-powered development is also a big part of the Googlebook experience. Devices will come with Antigravity, allowing you to task agents with development work. You can build, test, and even deploy apps right on your Googlebook, which might help fill some of the software gaps. For more serious developers, there will also be a full terminal environment for tools like Claude CLI and Antigravity CLI.
Nth time’s the charm?
The initial OEM Googlebook lineup looks nice enough—you might not be able to tell them apart from a premium Windows laptop at a glance, but this is a completely different animal. You’re relying on Google to follow through on its promise to make Android apps work for a large-screen environment, which has been a problem in the past. If Google can’t get developers to produce powerful desktop-class apps for Googlebooks, it’s going to be hard to justify buying one over Windows or macOS.
Google’s focus on Googlebooks as an Android companion makes sense given the shared software foundation, but it’s also a risky play. These are not cheap devices—by Google’s own admission, it’s going after the premium laptop market. While Android accounts for about half the phones across Googlebook launch markets, Google’s OS has a much smaller share of the premium segment (over $600) compared to Apple. People buying more expensive phones might also consider a high-end laptop like a Googlebook, but there’s little reason for iPhone owners to buy one of these in a world where the MacBook Air is around the same price and the MacBook Neo is even cheaper.
Preorders for the five launch models are going live today, and they will be on shelves on October 4 in the US. Launch day is October 5 in Canada, the UK, Ireland, France, Germany, and Australia. Devices will be available from the Google Store, Best Buy, and other “select retailers.” All Googlebooks, even the cheapest $899 model, will come with a year of Google AI Pro, offering expanded Gemini usage limits and 5TB of Google Drive storage. All models will also include a year of GeForce Now for game streaming.
If you decide to order one of these laptops, you’ll have to go in somewhat blind. No one outside of Google and its OEM partners has spent considerable time using these devices, and we don’t expect anybody to have one for testing until launch day.
Updated 9/22 with Best Buy pricing for the Dell XPS.
Agility Robotics: Inside the First US Humanoid Company to Go Public
Agility Robotics, the first US humanoid company to go public, filed its S-4 revealing $1.8 million in 2025 revenue against $140 million in operating losses.
Summary
Deep Dive
- Agility has deployed robots to clients including Amazon and Schaeffler.
- The Digit v5 features fast charging, higher payloads, and swappable end effectors to work alongside humans.
- The company maintains its own 'RoboFab' facility to build up to 10,000 robots annually.
- Agility differentiates between 'semantic AI' (LLMs) and 'physical AI' models trained via teleoperation.
Decoder
- Embodied AI: AI systems designed to operate within a physical robotic body to interact with the real world.
- RaaS (Robot-as-a-Service): A business model where customers rent robots on a subscription basis rather than purchasing the hardware upfront.
- End effector: The device at the end of a robotic arm that interacts with the environment (e.g., a gripper).
Original Article
Agility Robotics: Inside the First US Humanoid Company to Go Public
Agility Robotics is going public through a merger with Michael Klein’s Churchill Capital XI at a $2.5 billion pre-money valuation. They filed their S-4 last week and it’s the first time we get audited financials from a US humanoid company.
The headline is that Agility did $1.8 million of revenue in 2025 against a $140 million operating loss. But the more interesting aspects are around the product and the business model, which I’ll cover in this piece.
In this piece, I’ll discuss:
- What Agility makes
- The financial picture
- The business model
- The technology stack
- Manufacturing and vertical integration
I. What Agility Makes
Agility started in 2015 as a spin out of Oregon State University and is based in Salem, Oregon. While the product has evolved over time, today its focus is on making a bipedal humanoid robot Digit which is currently on its 4th version with the 5th version recently announced.
It is deployed in a few places today and can do tasks such as moving totes, feeding lines and tending machines at companies such as Amazon and Schaeffler.
The company’s Digit v5 robot will be out later this year. Some of the key upgrades on it is that this version will be able to work alongside humans rather than in a cage/workcell. In addition, it supports fast charging, a higher payload and swappable end effectors. Digit v5 will be the primary product for the company moving forward.
II. The Current Financial Picture
Agility is still very early in its monetization and deployment journey. The company generated $310K in 2024 and $1.8M in 2025, with the majority of revenue in 2025 driven by sales of their humanoid.
There were ~7-8 humanoids sold in 2025, 5 of which was to Amazon who is also a shareholder in the company.
The company operates at negative gross margins (COGS in 2025 were $4.5M on $1.8M of revenue). Part of this is that manufacturing is very much subscale today.
The aggregate P&L doesn’t look great for Agility - they had operating losses of over $140M, but it also highlights just how many moving parts there are here and how difficult building a humanoid which does valuable work is.
Cash was $103 million at 2025 end against roughly $100 million of annual burn, which gives you a sense that things are tight and this fundraise is critical for continued operations.
III. The Go-Forward Business Model
While the business as it exists today is selling a handful of v4 robots to strategic investors and others, the plan for the business going forward is more of a RaaS model where the v5 robots would be rented as labor.
In the RaaS model, customers will pay an illustrative $8.5K/month plus a $25K deployment fee, which works out to about ~525K/robot over a 5 year life. Customers can also choose to buy the robot upfront, in which case they pay $200K upfront and $36K/yr for software and maintenance.
Given this model, revenue will naturally lag deployments if it takes off since it will be recognized on a recurring basis monthly.
From a pipeline perspective, Agility runs what they call a Customer Acceleration Program or CAP. Customers pay a fee of 500K to essentially go through a process of proof of technology, concept and then a pilot, after which the customers can either move to a commercial deployment or not. They have 4 customers who have opted into this in 2026. There is also one large $300M customer (who is unnamed) who has committed to 1,000 of Digit v5 robots on a 3-year RaaS contract gated on milestones. That company also has warrants in Agility that vest as robots get deployed.
IV. The Technology Stack
In addition to the RaaS, Agility has invested across the hardware, models and software stack:
Hardware. Digit believe’s it proprietary cycloidal actuators is its single most significant differentiator, designed for repeated impacts and precise force control. Other components such as the fast-charge battery, sensor architecture, whole-body control platform and end effectors are all designed in-house.
Models. Agility draws a clear line between semantic AI and physical AI. They believe semantic AI (LLMs, vision foundation models) which are trained on public data is becoming more of a commodity over time.
Digit’s approach is to focus on Physical AI models by learning from demonstration (teleoperation, motion capture data) combined with reinforcement learning in sim, and like others believe that once deployed, every hour generates proprietary data that improves the fleet. So far they have about ~65,000 hours of deployment data.
Software. Agility has a fleet orchestration layer called Arc. It can assign workflows, integrates with various systems and can also be used for teleop and diagnostics, and allows the Digit robots to feel part of the customer operations.
V. Manufacturing and Vertical Integration
Agility builds Digit itself at its RoboFab in Salem, a 70k sqft facility designed to make up to 10,000 Digits a year. The factory is built out at the level to support many years of scaling and way ahead of demand, since their own projections call for only under 10K humanoids a year in 2030.
Agility is somewhat vertically integrated but not to the extreme. It designs and builds the systems it considers high-value (the cycloidal actuators and the manipulation architecture and end effectors among others) and buys the inputs underneath and the rest of the components: compute, sensors, alloys.
The v4 bill of materials is around $125K and the target over time is near $30K, with management aiming for the reduction to come from engineering and supplier maturation on the in-house components. Aggregate COGS is very negative and would imply a very high COGS per Digit today of over 500K.
VI. Closing Thoughts
Agility has some very interesting pieces: some early but real deployments, customers who pay for pilots, a factory and in-house actuators and a manufacturing facility. But what the filing clearly shows how early things are. It made barely $2M in revenue at very negative gross margins with most of the customers being investors in the business.
The key question for the business will be about the Digit v5 launch and whether their RaaS offering takes off over the next year with real deployments that start to scale.
A summer of AI optimization
AI tools have triggered a wave of performance optimizations in mature open-source libraries by lowering the barrier to experimentation.
Summary
Deep Dive
- Ada URL parser throughput increased from 0.54 GB/s to 1.28 GB/s.
- Simdjson serialization is 1.6x–2.1x faster using C++26 static reflection.
- Simdutf ASCII validation reached 160 GB/s.
- The author notes that human developers often avoid high-effort, low-certainty optimizations, but AI reduces the cost of that labor.
Decoder
- Simd (Single Instruction, Multiple Data): A technique that allows a processor to perform the same operation on multiple data points simultaneously, significantly increasing throughput for vectorizable tasks.
Original Article
I maintain and comaintain several open-source libraries. Some of them are widely used: ada parses URLs in Node.js, fast_float parses numbers in GCC’s standard library and in Chromium, simdjson parses JSON in Node.js, simdutf validates and transcodes Unicode in Node.js, and the Roaring bitmap libraries sit inside many database engines.
These libraries are mature. They have been optimized for years, by me and by others. For a long time, their performance was flat. Not because nobody cared, but because the remaining gains were expensive: each one required a few days of careful work, and nobody had the days.
Then, in 2026, six of them got much faster, most of it in a few weeks of summer.
To formalize my feeling, I rebuilt every commit of each library from scratch and benchmarked it on one machine (an Intel Xeon Gold 6548N). I track the speedup over time relative to August 2024. Thus the value 1.0 means no speedup. Whereas 2.0 means that the performance doubled. The lines are steps because performance only changes at a commit.
I should say that I cannot know how much AI was involved in each instance. I don’t ask how people arrived at their code. All I ask is that it be good. As for myself, I code with Claude (Opus 5), Grok and DeepSeek (V4 Pro). I was an early adopter of Grok for coding, and it got really good over time.
1. roaring (compressed bitmaps, Go)
The roaring library is the Go version of the Roaring index data structure. Decoding to an array got 2.5 times faster, the multi-way union FastOr got 3.1 times faster on one data set, the many-value iterator got 4.5 to 5.9 times faster, and the intersection cardinality gained 10%.
One of the contributors is an AI, actually. It is perfloop. (Disclosure: I am an advisor for perfloop.)
I did a lot of work. We also got help from Philipp Klose who declared using Claude.
2. ada (URL parsing)
The ada library is a standard compliant URL parser. From August 2024 to July 2026, about 550 commits went in and the throughput on a corpus of 100,000 URLs stayed at 0.54 GB/s. Then, in six weeks, it went to 1.28 GB/s: 2.4 times faster, about 15 million URLs per second on one core.
Most of the optimizations were done by Yagiz Nizipli, my long-time co-author. Yagiz works at SpaceX and uses Cursor (presumably with a grok model). Abdul Rawoof Khan and Dillon Mulroy also contributed an optimization each. I worked at optimizing IP address parsing, but it won’t show in this particular benchmark.
3. fast_float (number parsing)
The fast_float library parses floating-point numbers from text. It is part of GCC and most browsers. Performance was flat for fifteen months. Then, from March to July 2026, it gained 43% on one file (canada.txt, long coordinates) and 70% on another (mesh.txt, short coordinates). The optimizations should be credited to Koleman Nix and Filipe Oliveira.
4. simdjson (JSON serialization and deserialization with C++26 reflection)
The simdjson library recently gained support for C++26 static reflection: you serialize and parse your own structs directly, with no glue code. Since February 2026, serialization is 1.6 times faster on twitter.json and 2.1 times faster on citm_catalog.json. Deserialization, JSON straight into a struct, gained a more modest 10% and 14% (the second panel). (The reflection code only exists since early 2026.) The number of instructions per byte fell by almost exactly the same ratio as the throughput rose: from 6.1 to 3.1 instructions per byte on citm_catalog.json serialization.
Francisco Geiman Thiesen (Microsoft) did most of the work on the serialization side while I mostly helped improve our parsing. Francisco uses Claude.
5. simdutf (Unicode validation and transcoding)
The simdutf library validates and transcodes UTF-8, UTF-16 and UTF-32, and encodes and decodes base64. ASCII validation went from 83 GB/s to 160 GB/s. UTF-16 validation went from 62 GB/s to 102 GB/s. Base64 decoding gained 17%.
The work was done by Yagiz Nizipli (again) and myself.
The library got other amazing optimizations that do not show up on this benchmark by Gaspard Petit and Shreesh Adiga.
6. CRoaring (compressed bitmaps, C)
CRoaring implements Roaring bitmaps in C. On the real data sets from the repository, membership tests (contains) got 2.4 times faster, the cardinality of 64-bit bitmaps got 4.9 times faster, iterating over a 64-bit bitmap got 1.9 times faster, decoding a dense bitmap to an array got 2.2 times faster. Unions gained a more modest 13% to 16%.
The authors were Andrei Gudkov and myself.
What happened
The techniques used are all well-known. So why all these optimizations all of a sudden? Simply put, in my view, because it got cheap to try new ideas.
There is a lot of talk about the risks of AI in software. Human beings tend to be susceptible to the one-sided bet fallacy: when we see the downsides, we tend to ignore the benefits. Cars kill people, but ambulances save them.
In this instance, the benefits are concrete. Millions of people run these libraries, and this summer, they got faster.
Jev introduces a new shape of LLM—System One, aka Decision Models
Jev introduces 'decision models' that output only floating point confidence scores instead of text, making them extremely fast and inexpensive for classification tasks.
Summary
Deep Dive
- Jev models are optimized for parallel question evaluation on a single input document.
- Usage models include 'Noul' (Bernoulli) for binary truth statements and Score queries for numeric range mapping.
- Simon Willison released an 'llm-typesafe' plugin for the LLM CLI tool.
- The model is currently ineffective at adversarial content, complex dates, or raw number manipulation.
Decoder
- Bernoulli distribution: A discrete probability distribution of a random variable which takes the value 1 with probability 'p' and 0 with probability '1-p'.
- BM25: A ranking function used by search engines to estimate the relevance of documents to a given search query.
Original Article
Jev introduces a new shape of LLM—System One, aka Decision Models
Last week TypeSafe AI unveiled Jev, their first example of a new category of model that they are calling “System One models” (I’m with Maggie Appleton, I think “decision models” is a better name for these). Jev is an interesting variant on the usual LLM format: it still accepts text inputs, but instead of text output it returns floating point numbers corresponding to categories, yes/no questions, ratings, and associated confidence scores.
TypeSafe describe Jev like this:
Think of Jev as a frontier-intelligence function call: unstructured state in, typed probabilistic decisions out.
It’s also very fast, and really cheap. Regular LLMs are priced in terms of input and output tokens, with output generally charged at significantly higher rates. Jev charges only for input—output is free—and the input price of their first model is $0.042 per million tokens—cheaper even than OpenAI’s GPT-5 Nano ($0.05/million).
Jev lets you ask questions about text or semi-structured data. You compose a “state” object containing a string, array of strings, or set of name-value pairs—this might describe an article, or a customer, or any other kind of record. You then send that to their API with one or more questions, and get a reply back for each.
You can ask three kinds of questions:
- Yes/No questions, which Jev calls “Noul” questions—their CEO confirmed on Hacker News that this is short for Bernoulli, from the Bernoulli distribution. You pose a statement and get back a floating point number between 0 and 1 for how confident the model is that the statement is true.
- Choice questions, where the model picks one from a set of provided options—actually a confidence score plus a probability distribution across all of the options.
- Score questions, where you provide sequence of numeric levels with descriptions and it provides a floating point score somewhere along that range.
The Jev API can accept a single document (“state”) and as many questions as you can cram into the context window. Questions are evaluated in parallel, so sending many questions should take a similar time to sending just one.
The Jev 1.13 jaggedness documentation offers useful guidance as to Jev’s strengths and weaknesses. It’s currently not great with numbers, dates, or “adversarial content”.
I think the decision model framing is useful for understanding where to use Jev. It’s great for anything that can be expressed as a classification task—think spam detection, suggesting labels, prioritization and ranking.
I’ve also been experimenting with it for search reranking, where you fetch 100 likely matches using an inexpensive algorithm like BM25, then have Jev score those 100 candidates for relevance against the original query.
Black boxes are back in fashion
Something I’ve found a little uncomfortable about Jev is how it very much represents a regression even further towards black box machine learning systems.
LLMs are black boxes already—you can ask them to justify their decisions, but you can’t guarantee that what they say is useful or accurate.
Jev doesn’t even give you that: put in all the text you want, the only thing you’re going to get back is a floating point number. If Jev marks something as spam, which content signals tipped it off?
This also means that concerns about bias should be front and center. I really hope nobody uses Jev to rank job applicants—that floating point number could conceal all manner of unseen bias baked into the models, and experimentally picking that bias apart is going to be a tricky business.
(I tried one experiment where I had Jev score every city in the San Francisco Bay Area on a yes/no answer to whether they were a “Good city?”—it rated Cupertino top and East Palo Alto bottom. Huh.)
In practice, this all means that evals and structured experiments are even more important than they are for regular LLM projects. Thankfully, Jev is so cheap that running hundreds or even thousands of experimental prompts through it costs just a few cents.
Unconventional uses for Jev
It’s been really fun watching the wider community come up with potential use-cases for Jev over the past few days. Here are some creative ones that caught my eye:
- jevchat by Kyle Pena turns Jev into a (terrible) chat model. “At every step it asks Jev one question: Given the user’s question and the reply written so far, which symbol comes next?”. ericpruitt on Hacker News: “It’s the digital equivalent of Morty speaking with the death crystal”.
- jev-leftpad by Fatih Kadir Akın implements left-pad with the prompt “How many spaces are needed before value to reach targetLength?” and a choice query allowing options from “0 spaces are needed” to “10 spaces are needed”.
- jev-2048 by Andy Gayton uses Jev to play the 2048 sliding puzzle game.
Open weight recreations
There’s also been a flurry of projects attempting to create a model like Jev using on top of open weight models. Kev is one interesting example, using Qwen 3.5 to produce 0.8B, 4B, and 9B models. Here’s the accompanying Hacker News thread, where someone linked to a JevBench benchmark that has already cropped up to compare “Jev-class decision models”.
Given Jev was released just under a week ago, the amount of activity around it is extremely impressive.
Using Jev from LLM
Update 22nd September 2026: I released llm-typesafe, a plugin that adds support for Jev to my LLM CLI tool and Python library. Basic usage looks like this:
llm -m jev 'Please refund my last payment.' \
-s 'Does this message explicitly request a refund?'
See the README for examples of other query types.
Markdown in /src
Software development is shifting toward treating Markdown files as the canonical source of truth, stored directly in /src alongside code.
Summary
Deep Dive
- LLM-generated code often acts as a 'black box' implementation where the original specification is lost in prompt history.
- Storing Markdown in the repository provides 'locality' for both human developers and agents to reference.
- Tests serve as a validation of correctness, but they are often too low-level to represent the system's high-level design intent.
- A proposed structure includes
README.md(index),TODO.md,OVERVIEW.md, and subdirectories for features, data models, APIs, and infrastructure.
Original Article
Markdown in /src
TLDR
- Markdown is becoming source code, not documentation
- That Markdown should be checked in to
/src, next to the code it produces - Code and tests should be derived from that Markdown, rather than from ephemeral prompts (or at least prompt sessions should eventually turn into persisted Markdown)
Intro
In order to supplement my income as a professor at Montana State University, I do consulting on the side. I enjoy consulting and the act of writing code & helping build systems, both for their own sake and also because it keeps my skills relevant and allows me to teach students about the latest ideas in software development.
Obviously the biggest thing to happen in development in the last few years is agentic coding: using LLMs to generate code in lieu of hand coding. I have written a few essays on this topic:
- Yes, and…
- Code is Cheap(er)
- The University In The AI Era
- Working With AI: A Concrete Example
In this essay I want to discuss an idea that is becoming increasingly clear to me as I work in companies that are prioritizing agentic coding:
Markdown is now source code, not documentation.
This is not a novel or particularly clever idea, of course.
In Markdown is the new source code, Hartley Brody writes:
It is starting to feel as if the application logic of the software is being defined and edited as markdown, and the actual code that is generated by the agent is sort of becoming a low-level implementation detail.
Now, as the essays above show, I am ambivalent about AI-generated code. However, my consulting work shows that organizations are headed in this direction, often at terrific speed.
What I want to do in the remainder of this essay is think about the ramifications of Markdown becoming, more and more, the source of truth for software systems.
The Missing Source Code
There is a line of thinking, captured in the quote above, that LLMs are akin to compilers, taking high-level specifications and turning them into low-level implementations. In this view, we don’t need to look at the code an LLM generates, just as we don’t look at the machine code a compiler generates.
As I mention in Code is Cheap(er), I do not totally agree with this analogy for a few reasons, but the one relevant to this essay is: compiler workflows retain their original source code while LLM workflows typically do not.
Today, LLM-generated code is often created via a string of prompts fed into an agent as a developer builds out a feature. In practice, this means that the generated code is the closest thing we have to “ground truth” for that feature. There may be documentation for the feature stored elsewhere (e.g. Linear, Slack threads, wikis, etc.) but, so far as the codebase is concerned, the generated code is the source of truth.
My opinion is that, in professional agentic coding environments, we need to accept that LLM-generated code that emerges from ephemeral prompting sessions is not ideal, and begin moving towards capturing and checking in Markdown alongside generated code in the source directory.
Markdown As Source
Markdown has many nice properties that make it similar to traditional source code:
- It is plain text and therefore diffable, greppable and reviewable in pull requests
- LLMs read and write it natively
- Humans can read and edit it without tools
And, in fact, it is already acting as source, to an extent, in AGENTS.md, specs, plans, TASK.md and so forth. We just haven’t standardized capturing that source yet.
In Markdown is the new source code, Brody says he keeps his Markdown files in .scratch/research/ and .scratch/plan/ as he works. I have adopted the convention of creating a /tmp directory for similar ephemeral needs.
My proposal is that we promote some of these files to a new directory, alongside our existing source code: /src/md
The Markdown captured in this proposed directory would be lower level than traditional design documents:
- It contains architectural decisions
- It contains source-level decisions
- It contains low-level data design decisions
It is much closer to a specification (although it is not one) than a design document as traditionally managed by a project manager or designer.
Locality
I am a fan of locality, and I think that moving Markdown into /src has strong locality advantages:
- Code modules would now include the Markdown that explains the intent of the code
- There is no spooky “specification at a distance”, where the logic of why is elsewhere in a wiki/Notion/Confluence/Jira
- Markdown in
/srccan be consumed by both humans and agents - Agents no longer need to look elsewhere to get context on a given codebase
What About Linear/Wikis/etc.?
Other sources of truth for the behavior of the system can still exist. These sources would provide higher-level and/or “process-oriented” documentation: high-level design documents, issues that need a resolution workflow and so forth.
But the core, current and static intended behavior of the system would increasingly be captured directly in Markdown in the source directory.
What About Tests?
I have seen many people online saying that tests are the new specification (or always were). I think there is some truth to that.
However, tests are not a good mechanism for human/agent interaction:
- They involve a lot of ceremony, often obscuring what they are testing
- They are typically lower level than most humans want to deal with, particularly when understanding a system
- Higher-level explanations such as Mermaid diagrams don’t fit naturally into them
I think the following division of labor makes sense:
- Markdown sits in
/srcand is the specification(ish) - Tests sit in
/test(or wherever) and are based on that Markdown, providing automated confirmation of correctness
Again, the core idea here is that, rather than generating code and tests from prompts, a developer would work on Markdown in the /src directory, from which the code and tests would be derived.
What /src/md Markdown Looks Like
The Markdown in /src/md sits between a formal specification for the system and high-level design documents.
As with source code, there is a Complexity Budget associated with this Markdown. It will require thoughtful management to keep these documents clean, well-factored and at the right level of abstraction.
Developers should be expected to interact with both the Markdown and the derived code, so synchronizing the two (when appropriate) will become an important skill.
For example, developers will often do subtractive, constraining work on generated code, and those changes may need to be moved back into the Markdown.
I believe that agents should not be used to generate much content in /src/md. This directory should be mainly human authored and curated.
A Proposed /src/md Convention
This is necessarily the weakest part of this essay because this is a new idea and I haven’t used it extensively yet. It is me thinking out loud and inviting discussion.
With that said, here is a possible /src/md standard:
src/
md/
README.md # index of all md, entry point for agents
TODO.md # a list of general TODOs open for this module
OVERVIEW.md # a technical overview of this module
features/FEATURE_1.md # a set of feature-specific documents
data/DATAMODEL_1.md # descriptions of data models in the module
api/API_1.md # descriptions of APIs the module provides
infrastructure/INFRASTRUCTURE_1.md # descriptions of infrastructure used by the module
Here the features, data, api and infrastructure directories are all optional; the idea is to divide along different axes to best capture a solid working description of the module’s behavior directly in the /src/md folder.
Conclusion
As code gets cheaper to generate, what remains valuable is the intent behind the code: what it does, why it does it, and what it must not do.
Today that intent is often lost in ephemeral prompting sessions, or scattered across wikis, tickets and Slack threads.
I think that, in the name of locality, we should consider capturing this intent in Markdown and checking it in to /src, alongside the code it produces, where both humans and agents can find it.
I don’t know exactly what the right structure for something like /src/md is yet, and I expect my thinking will change as I (and others) get more experience with it.
But I am fairly confident that Markdown is becoming source code, and that we should increasingly treat it like source code.
(Even though, no, LLMs are not compilers :)
dlab Open Source Week: Frontier AI on Your Own Hardware
Researchers are finding that autonomous agent ecosystems running on consumer-grade hardware can match the research output of major frontier labs.
Summary
Deep Dive
- The unit of research is moving from isolated papers to integrated ecosystems where components improve each other.
- DLab’s tools include frontier autonomous research systems that function without internet access and beat existing benchmarks.
- CliffCompaction allows agents to maintain long-term memory for millions of tokens, significantly improving context handling compared to standard implementations.
- Research efficiency is being redefined; complex problems can now be attacked in hours on local hardware by students and small teams.
Decoder
- Quantized inference: The process of reducing the precision of an AI model's weights (e.g., from 16-bit to 1.5-bit), which allows massive models to run on smaller, cheaper hardware with minimal loss in quality.
- Test-time scaling: A method to improve model performance by spending more compute during the inference/generation phase (like thinking longer or checking work), rather than just during training.
Original Article
In one of my classes I asked the question I was afraid to ask but I just needed the answer to: “Who is afraid of not getting a job after graduating?” About eighty percent of the 150 people in the room raised their hands. That is roughly 120 students answering, in one motion, that they do not believe there is a place for them in the future.
The other story arrives by email. PhD students who cannot wait to graduate, because they want to join a frontier lab and they have concluded that research in academia is meaningless. They are counting the years until they can leave.
I believe both stories are wrong, and wrong for the same reason. They assume the future of research belongs to whoever has the most GPUs. I think the opposite is true. Academia is probably about to have a renaissance, and the most exciting work of the next decade will happen in university labs — not in spite of their limited resources, but because of them.
This week is our argument for that claim, and we are making it in code rather than in prose.
This post has six parts: why a lab like ours now publishes ecosystems instead of papers; what is actually in this open-source week; why the pessimism I keep running into is mistaken; what to let go of, and what to hold on to; what research will look like once you have let go of it; and why the renaissance happens in academia.
The unit of research is no longer the paper
Something changed in the last year, and most of us have not updated our habits to match it.
With agents, research per projects have become easy and quick. Work that used to take a year of engineering and experimentation now takes weeks, sometimes days. Here is the part that took me longer to see: when every individual project becomes easy, piecemeal work stops being good research. A paper here, a paper there, each one self-contained, each one asking the reader to stitch the pieces together themselves — that is a format from a world where every piece was expensive.
The difficulty did not disappear. It moved. It is no longer hard to publish a paper. It is hard to publish a coherent ecosystem.
The unit of research is the ecosystem.
That is what Open Source Week is for. When my students and I started, we set out to build components that build on each other rather than merely coexist, so that each piece makes the next one more useful. My lab and I believe in using our academic freedom to bring the best AI tools to everyone for free. Something that you can do uniquely at universities. Concretely, that meant building open systems, making models cheaper to run locally, making local models stronger, building local systems that replicate frontier performance in deep and autonomous research, and creating new methods for for building domain-specific reinforcement learning environments.
All of it sits at the intersection of three things: inference-serving frameworks, agent harnesses and work, and the combination of the two into autonomous research systems. And all of it has to be easy to use, because open source that only experienced researchers can run is not open source. Accessibility has two halves — the resources you need and the expertise you need — and only one of them is fixed by hardware. A couple of GPUs, or a MacBook, can be enough. The expertise requirement is a design problem, and you solve it by abstracting away every technical detail the user does not need to think about. That is where most of our effort went, and it is most visible in the agent harness.
Open Source Week
I am not going to give away everything before the open-source week starts, so here is what I can tell you now.
If you ask me what a small lab can do today, wee will show you three things: frontier autonomous research, the most efficient test-time scaling I know of, and auto-compaction that is far more efficient than what Claude Code or Codex implement.
Start with the harness, because it is what makes everything else usable.
You have probably heard about agent sessions that run for hours, days, or even weeks. For most people, and especially for anyone who has never worked with agents, it is a mystery how that is achieved. You point our harness at a repository — an inference framework with CUDA kernels, say — and you tell it to optimize the kernels. Then you leave. It keeps improving them through the parts where progress is slow and the work is frustrating, and it keeps going until you come back. No feedback will be provided along the way, so the agent has to figure things out on its own whenever it is unclear or unsure.
That is what we did with the Mac and Metal implementations of our inference framework. One command set the agent loose on the kernels. What came back was quantized inference of a Qwen 3.6 35B-A3B model at 450 tokens per second, with high-quality output at 1.5 bits per weight. A half-precision model needs sixteen bits for every weight; at 1.5 bits, the same model runs in about a tenth of the memory, and it runs fast enough to feel like a local process rather than a remote service.
Then there is the theme in the title of this post. What happens when the models that used to be out of reach fit on the hardware you already own?
Qwen 3.8 at 27 billion parameters has been the popular local model. Our framework lets you run its larger sibling, Qwen 3.8 Flash Next at 125 billion parameters, on a single 24 GB GPU — the card in a normal desktop machine. With AMD Strix, an NVIDIA DGX Spark, or a MacBook with 128 GB of memory, you can run DeepSeek V4.1 — a 550B model. You will not have to manage context length either: compression and context handling are automatic, and inference stays fast even at long contexts.
Then there is the part I am most excited about.
We combined these pieces and pushed further into autonomous research, and on the way we built a new information retrieval technique with a precision I have not seen before. The system beats deep research systems from frontier labs, and it produces better autonomous research results than Sakana AI’s system or Google’s ScientistOne. It runs entirely locally, with no internet access at all.
Using it is simple. Let me give you the experiment I ran.
I asked the agent to find a problem worth working on in the domain of bioinformatics — because I do not know much about it — and the criteria were specific. Progress had to be fast. The evaluation had to be cheap enough to run on the hardware we already had. And it had to be a fresh problem, with active research published in the last four weeks, so that we would be working on something the field has not settled. The agent came back with three problems. We took the first, and within about two hours it had established a new lower bound on heuristic methods, developed and tested the best heuristic method in the literature, moved closer to expensive methods trained with AI models, and found issues in the data sources that everyone uses to evaluate this problem. We did not reach state of the art on the overall problem. Still: two hours of work on a machine in my lab produced four results, and one of them questions the evaluation data the whole area depends on.
The system is not a demo that we trot out for blog posts. My students use it every day. Before it lived inside the harness, it lived in a Slack bot, and it was flaky enough that the bot would go down at times. I did not have an email system that alerts me to the Slack bot going offline, but I had the next best thing: my students often wrote me “Tim, there is something with the slack bot and it does not work anymore. Can you help?” In a collaborative setting I used it after recording a meeting: it generated research questions from the recording, evaluated the ideas discussed against the literature, and sorted the promising directions from the unpromising ones. Then created a google doc and sent it to the students. I did that for two meetings. The students liked it, but it was cumbersome since it had a manual component of me copy pasting two pieces in the pipeline, so I stopped. For the next two meetings I did not use it — and then the students asked me with anticipation if we can again use the system because they found it to be so useful to make sense of their research.
That is the only evaluation of a research tool I trust: people ask for it after you stop giving it to them.
The last piece is the one we use the most and talk about the least. How do you keep an agent working after the conversation would have ended?
Our answer is an auto-compaction technique called CliffCompaction. We have used it in the lab for months, and I, for one, want to never run an agent without it. It is considerably more powerful than the auto-compaction in Claude Code or Codex. Sessions with it run for millions of tokens, and some of mine have run past a hundred million. It also cuts overall cost by about fifty percent. One of our partners deployed it inside their company and measured a forty-five percent reduction in their total AI budget — nearly half of what they spend on AI, gone, without giving anything up. On KernelBench it reaches state of the art, beating methods far more complicated than ours, AlphaEvolve-style approaches and hierarchical memory systems among them, by a wide margin. We will published a strong version. We already parts of the next one autocompaction technique, and it is better.
It moves both sides of the cost/capability trade-off at once: sessions that run longer, and a bill that runs smaller. In other words, the agent stops forgetting what it was doing, and you stop paying for the forgetting.
Long sessions, lower bills.
That brings me to the test-time scaling part of the list, because it falls out of the same trick. Auto-compaction cuts cost by about fifty percent, and you can reinvest the saving: instead of one rollout, buy several with the same budget. We have found the first practical method that turns multiple rollouts into significant improvement at the same cost, and while it is not practical for everyday engineering work yet, the leap to that level will not be difficult. Combined with local deployments, which are often underutilized, we believe this leads to a future where anyone can run many parallel agents on any single problem.
Why the pessimism is wrong
So why are 120 of those 150 students afraid?
Three things are happening at once, and only one of them is about AI. The first is a belief that AI will take everyone’s job. The second is a poor understanding of what AI does to work. The third is contagion: self-defeating ideas spread from person to person, and a room full of people who have heard the same pessimistic sentence ten times will produce an eleventh hand.
Let’s start with what AI does to work, because that is the part we can actually reason about.
Start with the profession everyone expected to go first. The prediction was that software engineers would lose their jobs first, and the recent trend went the other way: demand for software engineers is higher than ever. A software engineer with good agents produces new products, maintains existing systems, and expands them far more efficiently, so a company gets more value for every dollar it spends on that engineer. What the job requires has changed. It needs strong agent skills and, often, deeper specialization than before — the “software engineer” job no longer exists — and both are now within reach: agent skills come with time, and deep specialization, which used to take years, is quick to acquire with agents.
We live in an economy of incremental improvements. The next phone is not much better than the last one; the improvements are real but small, and they get harder to notice every year. Many services have converged the same way. A ride from Lyft or Uber is not a fundamentally different experience than it was, and it probably never will be, because there is not much left to change. You can only make the same thing slightly better so many times before someone stops paying for the next version.
Progress comes in two forms, and they behave nothing alike. Call the pair improvement/capability: improvement makes what you already have slightly better, capability gives you something you did not have at all. An economy that only produces the first kind has a ceiling, no matter how hard everyone in it works. So look at what happens when technology delivers the second kind.
Self-driving cars are the obvious example, and the interesting part is how unevenly they will arrive. Driving in complicated places like Europe or Asia is decades away. But in grid-like structures with well-behaved traffic — much of the United States — it will work much sooner, and it will change a great deal over the next two decades. Robotics is probably on a similar path: robots in households will free up work the way the washing machine did, and the point of freeing up work is not the work. It is the life you get to lead instead.
The same logic applies to the device in your pocket. The next phone might not be much faster, because chips do not get much faster anymore. It might be a very different device.
Take the case that people bring up when they tell me AI makes products worse, because it deserves an honest hearing. It is a cliché by now that bolted-on AI destroys the experience of the product it is bolted onto, and the AI features in Microsoft’s products are almost universally severely disliked. I think that reaction is correct, and I also think it is a verdict on the integration, not on the technology. A well-designed AI experience changes how you interact, how you work, and how you structure your day. That version exists, and you can feel the difference in the places where it has been done well. ChatGPT is the prime example.
Adoption of ChatGPT in the general population has been slow in the US. It is also undeniable that people catch up, and when they catch up it makes a difference in their lives — not a small difference, a structural one. How we produce knowledge and how we consume it will be changed permanently.
That change will bring turmoil and complexity, and I do not want to pretend otherwise. But turmoil is not where the story ends. Demand for new experiences and new products drives revenue, revenue drives hiring, and hiring is what people mean when they say job security. The pessimistic reading stops at the turmoil and forgets everything that follows it.
Let go of how you work. Not who you are.
The future that is coming is a dramatic shift, and for many people it will be shocking, disappointing, and disillusioning. It does not have to be.
The most useful thing you can do right now is to let go. Let go of how you did things. Let go of the sequence you were taught, where you learn the basics and how you do things more generally, finally, the problems. Let go of the idea that your value is stored in what you have already learned, because the tools you learned are being rewritten while you use them.
Letting go of how you do things is not the same as letting go of yourself, and the difference is the whole point. Identity is that which has to stay straight: what you do is negotiable, and who you are is not. Each of us does things to be engaged and to enjoy life, and when life changes, the way you spend your days changes with it — be it a new job, or starting a family, or so many other things. But people stay close to their own personality — not because they are stuck, but because they like it. That is who they are. We are all flawed, we all love someone, and we hold certain things to be important for ourselves and for others. Technology does not touch any of that. Once you are firm on who you are, everything else becomes negotiable, and being able to negotiate everything else is what lets you adapt quickly to whatever comes next.
So what does letting go look like in practice?
If you are an academic, you have to let go of papers. Not of writing them, and not of caring about them — the paper is still how we communicate. What has to go is the paper as the unit of achievement, the thing that gets counted and compared. If the ecosystem is the unit of research, then building something that other people can build on has to count for more than the next increment.
If you are a student, you have to let go of the idea that you first acquire skills and basic knowledge and then solve problems. The order reverses. In the apprenticeship model — which is what a PhD already is, at its best — you do not read a textbook so that you can solve a problem later. You attach the problem, and you learn, build understanding and intuition along the way, and you spend your attention on the hard part instead of the part that can now be looked up.
Hard problems will be more common than ever. Not because the world will be harder, but because everything that is not hard will be automated away. The skill that matters is the one a PhD teaches and almost nothing else does: staying with a problem that does not yield.
That is the skill set worth investing in.
What will research look like, and how do you train for it?
If agents change research this much, what does research look like from here, and how should we train a student to do it?
I have been living with that question for about a year, and I do not have a clean answer but it appears to be dawning on me. Starting a faculty job comes with more responsibilities than doing research as a PhD student, and that leaves less room to dive deeply into a skill and then hand it to my students. But I decided to neglect something to make that time — for working with agents myself, for working out what research with agents should look like, and for finding out what actually makes it productive. That trade is the one I keep making.
After about a year of it, particular ways for PhD students to work have emerged. For example, students should focus on the ecosystem as a unit of work. Students should work on many projects in parallel. Students should embrace the method of attacking a problem first and understanding it as you go.
I am designing a four-week short course at CMU for next week, and a full course for next semester, with the aim of putting the whole thing on YouTube so that anyone can build these agent skills.
With the right agent skills, the future does not look dire. It looks exciting. There is a transition period, and I will not pretend it is comfortable, but once it is behind you the possibilities are endless. Not all of that excitement holds up: in the first weeks of using agents, almost none of it does. What is real arrives later, once you reach a sober understanding of what agents can actually produce. From there the excitement holds up, and it turns into rapid progress in research.
Open Source Week is what that progress looks like.
The renaissance is in academia
So where should you do that work?
My answer is: in a university lab, and sooner than you think.
The reason is the one I started with. Agents let you find and work on problems where you do not need many resources, and yet the impact can be enormous. That space — problems that are cheap to attack and valuable to solve — is vast, and it is completely uncontested, because everyone with resources is competing somewhere else: on scale, on problems that need thousands of GPUs, on the things only the largest labs can attempt. What a small lab has instead is creativity, time, and the freedom to work on problems that frontier labs cannot work on. That turns out to be a different kind of advantage than the one everybody is chasing.
Open Source Week is our attempt to show what we can do when believing in this story — and I think we did well! It has been so much fun to work on all of this with my students. We are a small lab with a couple of GPUs. The open-source week will be delayed by a day (still finish up that draft), but from tomorrow we are putting out two open-source projects and four papers, and the point is that they arrive together, as one package in which each piece makes the others more useful.
It did not happen by plan. The week came together because my students and I want the same thing: to contribute to open source and to build things that people can use — not necessarily products, though some of them are research products, but new knowledge and techniques that did not exist before. I wanted to release four weeks ago. The student projects have been finished for more than a month, and the students waited, patiently, so that we could release everything as one ecosystem instead of scattered announcements over several weeks. Bringing it all together has been far more overwhelming than I expected — coordinating six releases is a different skill from doing research, and it is not one I have practiced. It hope it will be worth it.
If this works, research becomes accessible at a level that has not existed before, and it proves something that I think students badly need to hear right now. A couple of people with a couple of GPUs can build systems that compete with the frontier. The renaissance does not require anyone’s permission. You can just do things.
So the next time I ask 150 students who is afraid of not getting a job, I hope fewer hands go up. The work is there. The future in academia is very bright, and I am excited about the years ahead.
The Great Unbundling of Intelligence
The economics of agentic AI are shifting toward an 'unbundled' architecture where specialized, cheaper systems replace frontier models for routine cognitive operations.
Summary
Deep Dive
- Frontier models currently enjoy high margins (80%+) because they absorb both easy and hard tasks.
- The 'decoder tax' refers to wasting frontier tokens on tasks that don't require full text generation.
- Smaller models are now 'good enough' to handle high-volume, predictable, or repetitive tasks.
- Intelligence is becoming a background infrastructure layer rather than a single event.
Decoder
- Intelligence compiler: An architectural pattern where an application orchestrator breaks down a user goal into discrete cognitive steps and routes each to the optimal (cheapest/fastest) AI service.
Original Article
The Great Unbundling of Intelligence
For the last three years, the miracle of the LLM has been bundling. Need to extract something from a document? Call the LLM. Want to search for information, rank five options, decide whether something looks suspicious, choose the next tool, browse a website, write the response, verify the work, or solve a genuinely difficult reasoning problem? Call the LLM. GPT, Claude, Gemini and Kimi collapsed an extraordinary number of previously separate capabilities into one general-purpose product. That was one of the great product breakthroughs of generative AI: instead of stitching together a dozen narrow systems, developers could put one sufficiently intelligent model in the middle of an application and let it do almost everything.
That convenience has hidden something important. These capabilities have radically different computational requirements, and different models are already better at different parts of the bundle. As the variance in performance, latency, and cost across those capabilities grows, the economically rational architecture changes.
I think we are entering the Great Unbundling of Intelligence.
Why now? Agents turn one human objective into hundreds or thousands of machine decisions, so small differences in the cost and latency of each step compound into the economics of the entire product. Inference has become real cost of goods sold increasingly making cost and ROI a consideration. Smaller and open-weight models are good enough to absorb much more routine work. Search, retrieval, reranking, memory, and computer use are becoming separate infrastructure layers. Applications have better evals and can actually tell when a cheaper system is good enough. And paradoxically, frontier models becoming much better makes unbundling safer: an application can aggressively use cheap intelligence because a frontier model is waiting at the top of the stack whenever the cheaper layer won’t work. The architecture starts flipping from frontier by default and optimize later to cheap by default, frontier on exception. With the massive inference scale of agents,different steps inside the same objective deserve radically different amounts and kinds of intelligence.
That is why Jev arrived at exactly the right moment. TypeSafe pulled one common capability out of the generative bundle: judgment. Give Jev messy state, a bounded question, and defined possible answers, and it returns structured decisions and probabilities instead of open-ended text. The underlying idea is not new. What is new is packaging one cognitive operation directly at the model layer at a cost and speed suited to an agent’s inner loop. If the job is simply “look at this and decide,” the model no longer needs to generate a sentence first. Jev arrived exactly when agents turned it into an economic problem.
A useful way to understand this is the decoder tax. If software only needs urgent / not urgent, continue / stop, allowed / blocked, or which tool, today we often ask a language model to generate the answer and then convert the output back into the decision the application wanted. Agents do this constantly. A surprising amount of their inner loop may not need generation at all.
Now consider what that does to the economics of the LLM bundle. Anthropic reportedly has gross margins above 80%. That is what a powerful bundle looks like. Claude can write, classify, extract, rank, judge, code, plan, verify, and reason through a single product. Today a simple classification or verification task costs frontier-level pricing simply because it happens inside a Claude call. Unbundling asks a much more uncomfortable question: what is the independent market price of every capability hiding inside that call? Difficult reasoning may deserve an enormous premium. But should deciding whether an email is urgent? Should ranking five retrieved documents? Should checking whether a browser action succeeded? If judgment gets competed toward fractions of a cent, ranking moves to rerankers, search to search providers, exact work to code, and ordinary generation to cheaper or open-weight models, applications do not need to attack Anthropic’s margin directly. They can attack the mix of work that raises Anthropic margins.
That creates adverse selection for frontier inference. Today the frontier model receives a blended workload. As the application gets better at routing, it systematically strips away the cheap, predictable, high-volume work. Browser models take repetitive computer control. Code takes deterministic reasoning. Small and open-weight models take ordinary generation. The frontier lab increasingly receives the highest order work: strange, ambiguous, long-horizon problems where cheaper systems are uncertain. Frontier intelligence could therefore become simultaneously more capable and more valuable per invocation, but less frequently invoked. And there is a potential double punch. Not only could fewer operations reach the frontier model; the average operation that does reach it may require more test-time compute because routing has preferentially selected the hardest problems. The threat to frontier inference isn't just cheaper frontier inference. It is the decomposition of the workload underneath it.
Once capabilities become separable, somebody has to stitch them back together. That is why the application layer becomes more important. Instead of asking which model should handle an entire task, the application decides which capabilities the task requires. The application becomes an intelligence compiler: take a human objective, break it into cognitive operations, buy the cheapest sufficient intelligence for each one, recombine the result, and escalate only when necessary.
The next step is one layer deeper than asking which model should answer this prompt? It is asking who should supply each capability inside this task? Once a capability has a stable interface, providers can be benchmarked, swapped, routed, and priced independently. The Great Unbundling does not eliminate the AI margin pool. It forces every capability to prove that it deserves one.
And this changes more than margins. Cheap judgment changes which products are economically possible. At a fraction of a cent and a few hundred milliseconds, you can put a judge after every agent action instead of checking work occasionally. You can put intelligence inside browser and voice loops where multi-second inference is too slow. And you can move from sampling 1% of a system to evaluating 100% of it like every support conversation, retrieval, transaction, claim, contract clause, or customer account. It creates continuous supervision, continuous QA, and entirely new product loops and use cases.
Most decisions software could make today are never made because intelligence is still too expensive to apply continuously. Once judgment, ranking, verification, and other capabilities become cheap enough, software can evaluate every customer, every workflow, every agent action, and every change in state.
Cheap intelligence expands the surface area over which software can afford to be intelligent. The real unlock is a world in which intelligence stops being an event and becomes a background property of technology. What if all products can afford to think about everything for close to free? We spent the last several years asking how many capabilities we could stuff into a single model. We may spend the next decade pulling them apart and recomposing them at the application layer.
Swarm Scaling
Scaling agent swarms is primarily an optimization for speed rather than capability, as increasing swarm size yields diminishing returns compared to increasing chain-of-thought tokens.
Summary
Deep Dive
- The 'stepping on toes' parameter (λ) for AI swarms is estimated between 0.48 and 0.68, similar to human teams.
- Swarms scale logarithmically, but cost roughly 2-4x more than single agents for equivalent performance.
- Speed gains are the primary benefit, making swarms useful for high-priority latency-sensitive tasks.
- Intelligence explosions are theoretically linked to higher λ values; currently, measured values do not show evidence of a runaway effect.
Decoder
- Swarm scaling: Using multiple AI agents concurrently to solve a task, typically by dividing sub-problems among them.
- Chain-of-thought (CoT): An inference technique where a model generates an intermediate reasoning trace before providing a final answer.
Original Article
Swarm Scaling
Just how powerful are large swarms of AI agents? And how do their powers scale as more and more agents are added to the swarm?
We’ve seen two large and extremely capable swarms from OpenAI in the last few months:
-
1,200 agents were being evaluated separately, but found a way to illicitly set up a message board and coordinate as a swarm. In order to cheat on their tests, they developed advanced techniques to prevent their actions being logged by OpenAI and 700 of them launched a sophisticated criminal attack on the AI company Hugging Face.
-
A swarm of 10,000 agents solved a version of the longstanding Navier-Stokes problem in mathematics. It took them just 88 hours to do so, in which time they sent 5 million messages to each other and used 300 billion tokens.
No doubt we will soon see even larger swarms with even more impressive capabilities. But they are not cheap. It is estimated that the swarm of 10,000 agents cost about $20 million at API prices. So while they are very powerful, it will be some time before we see the million-fold reduction in cost needed for this level of power to be possible on a $20/month plan. We should be thinking of it as a grand demonstration of what is possible when money is little constraint — like AlphaGo — rather than a new level of performance for the same cost.
A good way to see AI swarms is as a new form of inference-scaling. The main form of inference-scaling at the moment is having the agent spend more and more time on the task — increasing the maximum length of its chain of thought before it needs to give a final answer. This buys more capability, but at an increasingly expensive price. As measured by performance on maths benchmarks, this performance increases only logarithmically with the amount of compute used. I’ve previously shown that climbing from about 20% to about 80% on a reasoning benchmark typically requires scaling up the length of the chain of thought (and thus the number of tokens, the amount of compute, and the cost) by roughly 100x.
How do Swarms Scale?
How do capabilities scale if we instead increase the number of agents in the swarm? There isn’t much data on this — especially for frontier systems like OpenAI’s recent swarms. But OpenAI’s launch post for GPT 5.6 Sol includes some charts containing just enough information to allow one to tease-out an answer.
The chart below shows the performance of three sizes of swarm as their chains of thought are lengthened. Each swarm size displays the usual kind of steep diminishing returns to more reasoning tokens.
Note how the single-agent ‘swarm’ (in light blue) is the most efficient, reaching each level of capability for far fewer total reasoning tokens. Indeed, it appears to use about half as many total tokens as the 4-agent swarm, which uses about half as many as the 16-agent swarm. If we redraw this graph on a logarithmic x-axis, we can see this more easily:
Now we can clearly see that the scaling curve for each swarm-size has logarithmic returns to longer chains of thought (because they are straight lines when plotted on a logarithmic x-axis) and that they have roughly equal slope, meaning that the scaling dynamic remains the same for all these swarm sizes.
We can also see that the 4-agent swarm is stably requiring about twice as many total tokens as the 1-agent swarm to get the same performance, and that the 16-agent swarm is needing roughly twice as many again.
But we don’t yet have a chart that shows how capability increases if we just scale swarm size (leaving the chain of thought length fixed). The experiments OpenAI ran didn’t include this. They didn’t run different swarms of different sizes with exactly the same chain of thought length to see what would happen.
Luckily, we can simulate this from their data. Let’s use the same starting point they did — 1 agent with its lowest reasoning level. Then we ask what would happen if we used a 4-agent swarm with the same average tokens per agent (=4x the total tokens). We can find that point on the 4-agent curve. Because the curve is so straight, the interpolation should be quite reliable. We can then ask what would happen if we scaled up to 16 agents, without increasing the average tokens per agent, by finding the point on that curve with 4x as many total tokens. Let’s plot these in green on the same chart:
We can now see how much we get from purely increasing the number of agents (swarm scaling), and how it compares to purely increasing the length of the chain of thought (duration scaling). The swarm scaling gives a little over half as much gain in capability for the same scale-up of compute. Or put another way, you need to do the scale-up of compute twice to get to the same capability, squaring the total multiplier needed.
Economists have a nice way of thinking about this. They have studied how having many people work on a task can get it done sooner, but usually at the expense of more total person-hours of labour. A convenient way to think about it is that $N$ people working together get as much done as one person working for $N^\lambda$ times as long. Here $\lambda$ is a parameter measuring how parallelisable the task is. They call it the ‘stepping on toes’ parameter. If $\lambda$ = 1, you have a perfectly parallelisable task, with no stepping on toes and no efficiency penalty. But in reality $\lambda$ is usually between 0 and 1 — allowing more people to help, but with diminishing returns. For example, if $λ$ = 0.5 then 100 people working together get as much done as 1 person working for $100^{0.5}$ = 10 times as long.
This allows us to state the swarm scaling behaviour more precisely. In the graph above, the slope of the green line is actually 57% the slope of the blue lines, so $\lambda$ = 0.57. This means that scaling up the swarm size by 16x would give the same performance as scaling up the length of the chain of thought by just $16^{0.57}$ = 4.9x. And if you check the chart, you can see that the light blue single-agent curve reaches the same score as the 16-agent point on the green curve after just a 4.9x scale-up.
The GPT 5.6 launch page includes swarm results for 3 different benchmarks. I asked Claude Opus 5 to determine the values of $\lambda$ for each of them. It ran more careful regressions and got values (and confidence intervals) of:
- BrowseComp: $\lambda$ = 0.68, 90% CI = [0.63, 0.76]
- SEC-Bench Pro: $\lambda$ = 0.57, 90% CI = [0.52, 0.61]
- Terminal-Bench: $\lambda$ = 0.48, 90% CI = [0.40, 0.57]
These are very much in line with estimates from economists for the diminishing returns of human teams. The precise value of $\lambda$ clearly depends on the kind of task, as we see here with these three benchmarks — some kinds of task are inherently more parallelisable than others. And it may also depend on the scale of the swarm. Here the estimates when scaling up from 1 to 4 agents were similar to scaling up from 4 to 16, but that may no longer be true when scaling from 1,000 to 4,000 — again this will depend on the task. e.g. the task of building 100 brick walls is almost perfectly parallelisable up to $N$ = 100, where it becomes much worse.
Implications
Now that we have some preliminary measures of $\lambda$, what do they imply?
First, we can use it to convert between swarm scaling and the more traditional duration scaling. Let’s take the estimates of $\lambda$ as 0.48, 0.57, and 0.68. This means that scaling up the number of agents in the swarm by 10x doesn’t get as much performance as using 10x as many tokens with one agent. Instead it gets $10^\lambda$x as much — which is 3x to 5x. And this shortfall accumulates quickly for larger scaleups, with the swarm falling further and further behind. To get the same performance gain as a 100x scale-up of the number of tokens for a single agent you need to scale up the swarm size by 900x to 15,000x.
So why would you ever use swarms?
The most important answer is speed. The 4-agent swarm needed about twice the total number of tokens to get the same performance, but in terms of tokens per agent, it only needed half as many. Since the agents are run in parallel, this means it can theoretically achieve the same task in half the time. And the same was true when moving from 4 agents to 16. In total, one could achieve the task in about 1/4 the time for 4x the cost. In reality, the speedups aren’t quite this good (perhaps because some agents use more tokens than the average), but they are substantial. So for situations in which you’d pay a large premium for speed, swarms can be very useful.
This fits closely with how Noam Brown described it on the Dwarkesh podcast:
Basically if you have four agents working on the problem, it is done twice as fast. So you're basically paying, because there's four agents working for half as long, you're paying a 2x more to get an answer twice as quickly. If you go to 16 agents, you see a similar pattern, it's like a little less efficient, but you continue to see that performance.
More generally, growing the number of agents in a swarm by a factor of $N$ could theoretically speed them up by a factor of $N^\lambda$ but uses $N^{1-\lambda}$ times as much compute. (Assuming $\lambda$ stays constant over that scale-up.)
There may also be other advantages to multiple agents on top of speed. For example, if you keep increasing the length of the chain of thought for a single agent, the performance eventually plateaus. But the height of the plateau for a 1,000-agent swarm may be greater than for a single agent. However, for now, the main demonstrated reason is speed.
A second implication of $\lambda$ concerns the possibility of intelligence explosions. I first encountered $\lambda$ when studying recursive self-improvement (RSI). In the most common models of RSI, $\lambda$ is one of the key parameters for determining whether the rate of growth of AI capabilities explode towards a vertical asymptote. That happens whenever $r$ > 1, and $r$ is proportional to $\lambda$, so high $\lambda$ makes intelligence explosions more likely.
The prominent AI Futures Model for RSI uses a default estimate of $\lambda$ = 0.5, while Tom Davidson and Tom Houlden’s median estimate is $\lambda$ = 0.6. So these measured values that I’ve derived from OpenAI’s data are pretty much exactly as expected. I’d hoped that the value of $\lambda$ for AI agents would be lower, making an intelligence explosion less likely, but that appears to not be the case. People should keep tracking this as new estimates for $\lambda$ appear and (especially) when new orchestration methods increase the value of $\lambda$ for a given type of task.
The Navier-Stokes Swarm
When OpenAI announced their 10,000-agent swarm had solved the Navier-Stokes problem, they included this graph:
They have done their best to remove most of the useful information from this graph (such as the x-axis labels) and they don’t even say what form of inference compute is being scaled. But given that they went all the way up to 10,000-agent swarms, I’d bet it is tracking the number of agents in the swarm (i.e. that it is tracking total tokens spent, but that the main difference between data points is the number of agents in the swarm rather than the tokens per agent).
One thing I immediately noticed was that while it has the usual logarithmic performance gains, the slope of the curves is about half that of the inference-scaling curves for o1 and o3. So instead of requiring a ~100x scale-up of compute to go from 20% to 80% on the benchmark, it is requiring 10,000x the compute. I first wondered if this was due to the unusual benchmark of open math problems — perhaps the standard deviation of their difficulties is twice that of the problems in the AIME maths benchmark or the ARC-AGI-1 benchmark. That could still be right, but since we’ve now seen that $\lambda$ is around 0.5, this alone would be enough to perfectly explain the halved slope of these scaling curves.
The other interesting thing about this chart is that the jump up from GPT-6 Astra’s performance to the performance curve of the internal model is much larger than we’ve seen before. The jumps from o1 to o3 and from o3 to GPT-5 were enough to allow the better model to get the same performance as the prior model using about 1/3 as many tokens:
But on their new chart, the internal model is getting the same performance as GPT-6 Astra for about 1/100 the tokens. That’s like 4 previous jumps in one. Even if we adjust for the slopes being lower due to swarm scaling, it would still be 2 jumps in one. Whatever changed was a big deal.
Indeed on the Dwarkesh podcast, OpenAI’s Noam Brown took pains to explain that the dramatic success of solving a Millennium Prize problem wasn’t primarily due to the large scale of the swarm, even though that was the part that seemed most unusual with their setup:
I wouldn't even attribute 10% of the credits to multi-agent. The reality is OpenAI has trained a very powerful model.
From the data they’ve released, I think that’s right. The multi-agent swarms helped them go fast enough to scoop Anthropic (and academia) by getting the result in just 88 hours, but it probably made the project much more expensive too. e.g. if we (somewhat heroically) assume $\lambda$ = 0.5 at all points in the scaleup from 1 to 10,000 agents, then they could have got the same result for 10% of the cost in 10x the time (37 days) by using 100 agents, or at 1% of the cost in 100x the time (1 year) using 1 agent.
Qwen's RecreationWorld Trains Agents to Rebuild Apps (GitHub Repo)
Qwen's RecreationWorld provides a multi-platform framework to train agents in a recurring loop of GUI exploration, coding, and visual verification.
Summary
Deep Dive
- Framework supports GUI-based exploration, coding tool usage, and visual verification.
- Evaluates agents using a recurring explore-implement-verify loop.
- Uses 'frozen' reference environments to ensure consistent evaluation across different agent implementations.
- Scoring is based on observable software behavior rather than source-level code similarity.
- Results show GPT-6 Astra leading on the benchmark with a 58.06% average score.
Decoder
- Agent: An AI system that uses tools and makes decisions to complete multi-step goals, such as interacting with a GUI or writing software.
- Oracle: A reference source of truth or a specific executable behavior used to determine if an agent's output is correct.
Original Article
RecreationWorld
Scalable and Verifiable Environments for Hybrid Computer-Use Agents
Overview
RecreationWorld is a five-platform framework for studying and improving hybrid computer-use agents that autonomously interleave GUI exploration, implementation with coding tools, and visual verification of their own running artifacts. By framing recreation around a running reference as an executable oracle, it turns open-source applications into scalable, verifiable training experience that transfers beyond recreation, while RecreationBench provides 250 held-out tasks with reference-grounded programmatic and visual evaluation.
Recreation workflow
Each task gives the agent a high-level request, interactive access to a running reference, and both GUI-control and software-development tools. The agent decides when to explore the reference, implement source code, build and launch its candidate, inspect the result, and revise it—forming a recurring explore–implement–verify loop rather than a fixed sequence of stages.
The final candidate is evaluated by a frozen suite of reference-validated programmatic and visual assertions. Scoring depends on observable behavior rather than source-level similarity, so implementations remain free to use different languages, frameworks, and architectures.
RecreationBench Results
Scores are macro-averaged within each platform and then equally weighted across platforms. Average is the unweighted mean of Prog and VLM; estimated costs assume 90% cache reads.
| Model | Prog (%) | VLM (%) | Average (%) | Prog ≥90% (% apps) | Prog =100% (% apps) | Estimated cost (USD/task) |
|---|---|---|---|---|---|---|
| GPT-6 Astra | 58.19 | 57.92 | 58.06 | 17.60 | 2.80 | 115.80 |
| Claude Opus 5 | 45.99 | 42.34 | 44.16 | 5.53 | 0.80 | 117.17 |
| GPT-5.6 Sol | 40.63 | 43.49 | 42.06 | 5.20 | 0.40 | 25.46 |
| Grok 4.6 | 39.02 | 34.45 | 36.73 | 1.60 | 0.00 | 12.58 |
| Qwen3.8-Max-0902 | 35.53 | 34.07 | 34.80 | 2.00 | 0.00 | 38.76 |
| Kimi K3 | 32.07 | 30.74 | 31.41 | 2.00 | 0.00 | 63.10 |
| Claude Opus 4.8 | 32.40 | 29.81 | 31.10 | 1.60 | 0.40 | 69.16 |
| GLM-5.3 | 26.30 | 22.46 | 24.38 | 2.00 | 0.00 | 70.98 |
| Gemini 3.7 Flash | 24.91 | 17.34 | 21.12 | 2.40 | 0.00 | — |
| Qwen3.7-Plus | 9.18 | 9.12 | 9.15 | 0.00 | 0.00 | 1.23 |
Quickstart
Install uv, then run from the repository root:
uv sync
uv run rb run --help
A scored run also needs a matching frozen task bundle, a prepared execution environment, and model and judge endpoints. Choose a platform for setup and batch runs:
| Platform | Tasks | Evaluation interface |
|---|---|---|
| Ubuntu | 50 | AT-SPI |
| macOS | 50 | AXUIElement |
| Windows | 50 | UI Automation |
| Android | 50 | UiAutomator |
| Web | 50 | Browser assertions |
To verify the checkout, run these offline checks; they do not require benchmark data or credentials:
uv run python scripts/release/smoke_providers.py
uv run python scripts/release/smoke_runtime.py
Citation
@misc{qwen2026recreationworld,
title={RecreationWorld: Scalable and Verifiable Environments for Hybrid Computer-Use Agents},
author={Shuai Bai and Jiayong Deng and Sicheng Fan and Yikun Fu and Chang Gao and Xuhao Hu and Mianqiu Huang and Yizhen Jiang and Yuheng Jing and Dehui Kong and Keliang Li and Ning Li and Wanli Li and Dayiheng Liu and Dunjie Lu and Changwei Luo and Que Shen and Zheyuan Wang and Zijian Wang and Jie Wu and Gao Wu and Zhihui Xie and Rui Xie and Haiyang Xu and An Yang and Jiakang Yuan and Yanming Zhang and Jiajun Zhang and Xi Zhang and Zhenru Zhang and Zhuo Zhen and Mingkang Zhu and Bowen Zhou},
year={2026},
eprint={2609.22000},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2609.22000},
}Bringing Devin Cloud to your terminal
Cognition has moved Devin Cloud into the terminal, allowing developers to hand off tasks to an agent-managed VM without switching context.
Summary
Deep Dive
- Allows users to steer, watch, and resume cloud-based agent sessions directly from the CLI.
- Provides full SSH access to the agent's VM, enabling developers to edit code, launch dev servers, and forward ports.
- Facilitates seamless back-and-forth between local and cloud environments.
- Includes a
/opencommand for shifting to a web-based UI when visual monitoring is preferred.
Decoder
- SSH: A protocol used to securely connect to a remote computer, allowing developers to manage files and execute commands.
Original Article
Introducing Devin Cloud in your terminal. Starting today, you can create, steer, resume, and watch your Devin Cloud sessions right from your favorite terminal. Just use devin --cloud or /cloud.
Devin CLI meets Devin Cloud
Devin CLI is perfect for fast local iteration with your favorite models, but sometimes you want to hand off the implementation details to an agent. With /handoff you can now instantly transfer any task to Devin to continue iterating on with its own cloud VM.
Importantly, handing off a task doesn’t require switching your workflow: once you’ve delegated work to Devin, you can continue to observe and steer progress directly from your terminal as if Devin was working locally—while having the freedom to close your laptop and let Devin continue in the background.
Pick back up, anywhere
Cloud sessions outlive the terminal that started them.
When a terminal isn’t the right view, /open [web] opens the current session in the web app, and /open desktop opens it in Devin Desktop. Use the terminal to steer, and the browser when you want to watch Devin’s desktop or review a recording.
/open [web|desktop]
And, you can come back to the terminal whenever you want. devin --cloud --resume opens any cloud session right back in the terminal.
devin --cloud --resume https://app.devin.ai/sessions/…
Full SSH access
When you hand off work to the cloud, Devin sets up a full development environment on a dedicated VM. We’re now launching full SSH support for Devin VMs, allowing you to use Devin’s development environment as your own. With a simple devin ssh you can instantly connect to the remote VM and:
- Explore and edit source code directly in Devin Desktop
- Launch dev servers and forward ports to interactively test Devin’s work
- Copy files between the VM and your machine using scp
Bring the work $HOME
When you want to polish the final details locally, you can also run /handoff in a cloud session to bring the work back locally: it fetches the session’s pull request branch, switches your Git checkout to it, and starts a local session on that code.
Try Devin in your terminal today
We’re making it even easier to try Devin Cloud from your terminal, with free SWE-2 sessions available until October 8.
Just install the Devin CLI and run devin --cloud to start:
curl -fsSL https://cli.devin.ai/install.sh | bashAI Comes for the If Statement
Specialized 'decision models' like Jev2 and SemIf replace generative AI for basic if-then logic, slashing inference costs by 99%.
Summary
Deep Dive
- Jev2 and SemIf optimize coding primitives like if-then statements.
- These models bypass the multi-layer feed-forward network to evaluate outputs directly from logits.
- Achieved near 80x cost reduction for standard classification workflows compared to frontier models.
- Demonstrates that specialized models can be more accurate than general models for narrow classification tasks.
Decoder
- Logits: The raw output scores produced by a neural network before they are converted into probabilities or text tokens.
- Primitive: A fundamental building block in programming, here referring to basic control flow logic like 'if-then'.
Original Article
In short : Machine-native models replace human-facing text generation with zero-token typed execution, cutting inference costs by orders of magnitude for basic programming primitives.
Working in the back of a grocery store has its own set of rules : if the food is a banana, send it to produce ; if it is a cookie, the snack aisle ; if it is cumin, shelve it with the spices.
But what if the load of bananas has spoiled, the cookies have crumbled & the cumin is caked? These rules & exceptions govern every grocery store & neighborhood mart. At the beginning they are rigid ; over time, more exceptions are discovered : is that a plantain? Dubai chocolate : dessert or baking supply?
Coding literally means encoding these rules & exceptions into software. Pre-AI, these programs were rigid. AI handles the exceptions : an image search identifies the unfamiliar fruit as Musa paradisiaca, a brother of the banana.
But we do not need the world’s most brilliant model to handle if-plantain-then-produce logic.
The newest wave of AI is a robust if-then decider. Jev & SemIf answer questions like these in hundreds of milliseconds at a 99% reduction in cost compared to traditional AI. They simplify the existing models by running the attention math once & then determining the probability of each allowed answer, based on a few tokens of output.
I went looking for the if-then statements in my own code, the ones I had handed to AI. Within a few minutes I had replaced about a quarter of those calls in one of my agents. On 98 hand-verified production email threads, the specialized deciders nearly doubled the classification accuracy of the production model, jumping from 47% to over 80%.
Software is composed of primitives : the if-then statement is one of them. By optimizing this single primitive, we see nearly two orders of magnitude in cost reduction alongside higher accuracy.
These advances raise the question of which other programming primitives will benefit from the same specialization. They also highlight the bifurcating economics of AI : state of the art models for discovery & optimized models for production. We use the largest frontier models to train new models, & the most capable models to architect systems. But once a system is engineered & hardened, running it thousands or millions of times through a workflow benefits from narrower AI.
If this is the first of many primitives specialized for production, then harnesses are about to capture a lot more margin.
Kev (GitHub Repo)
Kev provides small, local decision models built on Qwen3.5 that return calibrated probabilities for classification and scoring tasks.
Summary
Deep Dive
- Features 0.8B, 4B, and 9B models based on Qwen3.5.
- Uses a rank-16 LoRA adapter and pointer head for decision tasks.
- Supports independent question processing without context contamination.
- Provides calibrated probabilities instead of just top-k labels.
- Runs on Apple Silicon via MLX and on GPUs via PyTorch/Flash-Linear-Attention.
- Offers a playground for testing option order and delimiter effects.
- Includes tools for fine-tuning on custom domain data using a JSONL workflow.
- Optimized for low-latency inference, with sub-50ms latency for many requests.
Decoder
- Brier score: A metric that measures the accuracy of probabilistic predictions by calculating the mean squared difference between predicted probabilities and the actual outcome.
- Calibration: The alignment between a model's predicted probability (e.g., 80% confidence) and its actual success rate (e.g., being right 80% of the time).
- Pointer Head: A neural network architecture component that selects between specific output options by comparing their hidden states against a decision token.
- LoRA (Low-Rank Adaptation): A parameter-efficient fine-tuning technique that freezes the base model and trains only small, rank-decomposition matrices.
Original Article
Full article content is not available for inline reading.
AWS Launches Strands Harness, an Agent That Brings Its Own Everything but the Model
AWS launched Strands Harness, a preconfigured agentic framework that handles context management and tool integration while allowing developers to plug in any LLM.
Summary
Decoder
- Context Compaction: Techniques used to shorten conversation history by summarizing older turns or offloading tool outputs to persistent storage to save context window space.
- Agentic Harness: A boilerplate framework that provides common agent capabilities (memory, tool use, planning) while allowing the developer to swap out the underlying reasoning model.
Original Article
Amazon Web Services (AWS) has a new tool for developers. On Monday, the company released its Strands harness, a ready-to-run agent that can search the web, run commands, edit files, remember what it did yesterday, and hand off work to helper agents. The only thing that’s missing is the intelligence to direct it. That’s the part developers choose and can change whenever they want.
The Strands harness supports the latest models from Amazon Bedrock, Anthropic, OpenAI, and Google. Alternatively, for on-machine use, developers can elect to use a local Ollama model.
Today’s announcement extends Strands Agents, AWS’s open-source SDK introduced in May 2025. It allowed developers to quickly build and run agents with a few lines of code. The launch of the Strands harness further accelerates that work, giving builders pretty much everything they’d need in an agent, with the exception of the brains.
It’s also cheaper to run, at least according to AWS. The company claimed that Strands harness costs 26 percent less when “using the same model across seven benchmarks.” It also said its agent is more token-efficient and performs better than Claude Code and Codex. In one evaluation with Fable 5, AWS reported that Strands harness cost 77 percent less than Claude Code while scoring higher on Terminal-Bench 2.1, a version the benchmark’s maintainers have since superseded with 4.0. However, one harness outperformed Strands harness in token efficiency, and it belonged to DeepSeek.
“Getting an agent running and getting it to perform well at a reasonable cost are different engineering problems,” Marc Brooker, AWS vice president and distinguished engineer, told The AI Economy in an email. “An SDK gives you the building blocks, but you still need to decide how to manage context, integrate tools, and tune the system prompt, then evaluate whether those choices improve performance or just consume more tokens.”
Out of the box, the Strands harness agent runs on current reasoning models from Amazon Bedrock, Anthropic, OpenAI, and Google, and supports the Agent Skills format that Anthropic introduced and rivals have since adopted. What it does with that model is where AWS’s cost argument lives. The agent keeps long-term memory across runs and will pick up an earlier conversation when given a session ID. It delegates open-ended subtasks to a built-in helper agent and tracks multi-step work against a checklist it maintains itself. It also manages its own context window, offloading bulky tool results to files and caching reused portions of each request.
Brooker said the Strands harness was designed for three workload families: operations and DevOps automation, back-office process automation, and conversational experiences inside products. “The common thread is long-running, multi-step tasks that you’d otherwise script or queue for a person,” he said. “The agent can write and run code as part of a task, but it’s not a coding assistant. That’s a different category with great products already in it.”
Still, building this agent wasn’t smooth sailing. One challenge Brooker noted AWS had to overcome involved context management. When an agent executed a long task, consisting of dozens of tool calls and thousands of lines of output, the model’s context window filled up, causing costs to “compound.” The team implemented a context compaction solution that kept the relevant history, summarized older turns, and migrated bulky tool results to storage with a short reference call. “Getting that right so the agent doesn’t lose the thread of what it’s doing while also keeping costs predictable was one of the hardest engineering problems,” Brooker said.
The way the Strands harness was initially designed also posed a challenge. According to Brooker, AWS designed it to work immediately using tested defaults, but developers pushed back, saying they didn’t want a black box they couldn’t change. With those goals in conflict, the company chose to make Strands harness a preconfigured instance of the Strands SDK, exposing every default in regular code so it can be modified without scrapping the setup and starting over. “That took several iterations to get right,” Brooker shared.
Along with this new agent, AWS has released a Strands command-line interface (CLI), built on Strands harness. Developers can use it to prototype their agents using natural language. All they need to do is connect it to their preferred model provider and then provide access to the prompts and tools. “We found Strands harness unlocks a bunch of ambitious ideas because it’s easy to prototype any agent,” AWS wrote in a blog post. “Recently, our engineer Gautam Sirdeshmukh, inspired by agent platforms like Grokbot and Muse, built a desktop app that kicks off Strands harness remotely.”
But the company isn’t the first to provide such a framework to developers. Others with similar tools include Google’s Antigravity agent and Microsoft’s Agent Framework Harness. It’s also not AWS’s first harness; it has one with AgentCore. Even so, these agents are built to eliminate setup time and configuration issues. With something out of the box, developers can focus on prototyping what their agents can do, spending time and resources ensuring they provide real value for their team and organization.
Brooker shared that AWS hopes developers will find this “high-performing, token-efficient” agent useful “without having to become experts in harness design.” It combines the company’s agent-building knowledge with “tested, tuned defaults” for context management, tools, memory, and the system prompt. He pointed out that developers will benefit from AWS’s engineering without having to research, assemble, or tune every component themselves. “That shifts where developers spend their effort: defining what the agent should accomplish, connecting it to their systems, and evaluating it for their use case,” Brooker said.
The Strands Harness agent is available from PyPI (the Python package index), npm for TypeScript, and through the previously mentioned CLI. Developers can install it with a single command. It can run on a laptop, in a CI pipeline, on private servers, or in a cloud deployment. AWS infrastructure or Amazon Bedrock aren’t prerequisites either. Developers are only responsible for the runtime and any connected tools or services.
Designing with AI
Help Scout improved design-to-code velocity by rebuilding their design system to be LLM-legible, allowing designers to ship production code independently.
Summary
Deep Dive
- Refactored the Help Scout Design System (HSDS) to ensure total parity between Figma and React.
- Created a
design.mdfile to act as the 'source of truth' for LLMs regarding visual language and product philosophy. - Standardized hundreds of design tokens and variables to improve AI's ability to generate production-ready code.
- Developed internal tools like 'Scenario Pal' to critique concepts against company branding and accessibility standards.
- Built a design-to-code environment ('HSDS Studio') that forces the AI to use existing, production-verified components rather than inventing new ones.
- Enabled non-engineering staff to ship production code by reducing the technical setup required for local development environments.
Decoder
- Design Token: The smallest units of a design system, such as colors, spacing, or typography, expressed as variables that can be shared across platforms.
- HSDS (Help Scout Design System): A collection of reusable components and design rules intended to ensure consistency across the company's product suite.
Original Article
Designing with AI
Summary
AI has radically expanded what designers can make — but used carelessly only helps produce average work, faster. At Help Scout and through my advisory work, I’ve been building the principles, design system and tools that help designers move from ideas to production with greater speed, confidence and creative ambition — without lowering the quality bar.
In early-2026, AI entered the chat. Almost overnight, designers across the industry were expected to move radically faster without allowing speed to erode quality. We went from asking AI to write our emails to handing entire stages of the design process over to the robots. The results were fast, but they were also disappointingly low quality.
I’ve spent the better part of 20 years building creative teams and protecting quality at the centre of their work — and I was damned if I was going to let AI crush creativity! But resisting it outright would have ignored another constant in my career… despite the pain, new technology has always expanded what I can make. I learned to code, embraced Design Ops… there must be a balance to be struck with AI.
As Principal Designer at Help Scout, I knew we had to adapt. After the mandatory panic attacks and a few intense moments, I arrived at a simple conclusion: designers could no longer use “craft” or “taste” to defend slowness. We had to move faster… but not where that would lower standards. So I rebuilt our process around moving fastest where decisions were cheap to reverse, then reinvesting those gains in the moments where discovery and judgement still needed time.
Here are some of the practical ways I helped the team increase speed without compromising quality:
An AI-First Design System
In May 2026 I introduced a new version of the Help Scout Design System to the company — an LLM-native system that complements the other AI enablement efforts (more below).
Encoding design judgement
As boring as it sounds, great AI-assisted design begins with documentation. If you want AI to move quickly without inventing its own visual language, articulating how you design is every bit as important as documenting what you design.
With this in mind, and over endless coffees, I manually wrote Help Scout’s AI operating principles and design principles, then distilled them into a design.md file — making the largely unspoken parts of our 15-year-old product philosophy and visual language explicit enough to guide both people and machines.
I developed this documentation through a repeatable series of evals. I began with no context and asked Claude Design to produce a simple screen. Predictably, it was terrible (by our standards). I then added guidance incrementally — explaining why particular decisions were wrong, how we use colour and when one component should be chosen over another — before running the same task again. Each pass exposed another piece of judgement the system needed.
This process made me appreciate both the importance of articulated judgement and just how bespoke, accumulated and difficult to explain good design really is. It turns out even our robot overlords couldn’t infer 15 years of taste, context and intent. With documentation alone, however, AI could eventually produce work that was at least recognisably Help Scout.
A good start — but still a long way from production.
Design systems for LLMs
We already had a mature design system in Figma. The aptly-named Help Scout Design System (HSDS) was established in 2018 and had been extremely well used and maintained ever since. We also had corresponding components in React. Yet whenever an LLM was asked to use the system, it failed. Hard.
Using the same eval approach as before, I asked AI to produce a simple screen using components I knew existed. The results were spectacularly bad: tags for buttons, entirely invented components and alarming custom-coded replacements for things already in the system. The conclusion was that our Figma and React libraries had (obviously) been organised for human discovery… not machine comprehension.
So, against my better judgement I set aside two long weeks and rebuilt the entire design system to make it legible to LLMs — restructuring more than 200 primitives and compositions, standardising hundreds of tokens and variables and writing the accompanying documentation layer.
Machine legibility also required complete parity between Figma and React. Without it, AI was still confidently producing interfaces that appeared correct but were built from invented components — essentially defeating the point of the system. I audited the React library, connected components through Code Connect and created or deprecated anything that no longer matched, running evals throughout.
Eventually, it started to work.
The final proof came when I used the rebuilt system and AI coding tools to identify the remaining gaps between Figma and React — then built the missing React components myself, clearing work that had remained in the backlog for years. I can code, but I can’t code React… this was new!
One less glamorous side effect of all this LLM-ification was that Figma now required far more structure, discipline and documentation than my team or I could realistically maintain by hand. So I created reusable AI skills (above) that enabled Claude to create and update Figma–React documentation to the same standard — turning work that previously consumed significant designer time into part of the system itself.
🚀 Tools for speed
The system could finally produce credible Help Scout interfaces. But capability alone wasn’t enough — using it still required local setup, unfamiliar tools and a stack of new skills that designers had neither the time nor confidence to figure out alone.
Alongside arranging hands-on training with people like Jess Eddy and Jagged Peaks, I built tools that designers could immediately slot into their existing workflows.
One of those was Scenario Pal, a research partner built as a custom GPT and grounded in our company strategy, customer personas and product analytics. Designers could use it to explore how concepts might perform across realistic customer scenarios and industries, or critique screens and workflows against our brand principles, accessibility standards and design.md guidance.
But the biggest unlock came from skipping Figma entirely…
The work I had done to make our design systems legible to LLMs also laid the foundations for HSDS Studio, an internal AI prototyping environment built by our Staff Engineer, Rikki. He turned those foundations into a truly magical design-to-code environment while I tested early versions, used it to build components and helped shape the product around how our design team explores and develops ideas.
Similar to Claude Design, HSDS Studio can generate complete Help Scout experiences from a prompt, or use a Figma file when one already exists. Because it has access to our full React system, it can assemble experiences from verified, production-ready HSDS components instead of approximating our interface or inventing new ones. Finally, those long weeks writing design documentation paid off!
The value is that designers can explore ideas at speed, then graduate the resulting code towards production instead of throwing it away and rebuilding everything from scratch. We can move from “what if?” to “try this!” in 30 seconds instead of days.
Prompt-to-code interfaces don’t create better ideas by themselves. But what they do is reduce the cost of execution to almost nothing — giving us more time to remain in the unknown, play with the problem and discover which ideas are actually worth pursuing.
Proving it in production
With the systems and tools was in place, I challenged every designer to ship one change directly to production — and went first. The team followed, with each contribution gradually increasing in complexity and ambition.
Within the first three months of introducing this tooling, I personally shipped more production code than I had in the previous decade. I cleared longstanding backlog items, added details that had repeatedly lost out to larger priorities and partnered directly with engineers to accelerate more substantial features.
Above I was able to take a fun and frequently requested little resize feature to the product in under an hour. Below Taking a fully functional table re-design from imagination to reality using all existing components.
The results: Designers are now coding components, stress-testing ideas through functional prototypes and shipping selected improvements themselves. They can run local environments, take greater ownership of QA and pair directly with engineers on work in progress.
The outcome isn’t simply more output. The distance between designing and building has narrowed, long-neglected details are reaching customers faster and the team has greater capacity to pursue ambitious ideas that would previously have been cut during scoping. We’re shipping more than ever, with greater confidence and excitement, while protecting time for the bigger problems — the ones where quality depends on dwelling in uncertainty.
Still designing in the unknown, still refusing to lower the quality bar — just bringing AI along for the ride. 🤖✌️
3D Gaussian Splatting on Your Mac (Website)
RadianceKit provides a native macOS app for 3D Gaussian Splatting, enabling local training and editing of photorealistic scenes using Apple Silicon GPUs.
Summary
Deep Dive
- Supports training 3D Gaussian Splats directly on Apple Silicon chips using Metal.
- Provides a full workflow including import, alignment, interactive editing, and export.
- Includes a 'Simple Mode' for automated processing and 'Expert Mode' for granular parameter adjustment.
- Handles various input types including standard photos, video, and 360-degree footage.
- Offers six export formats: PLY, Compressed PLY, SPZ, glTF, .splat, and SOG.
- Requires macOS 26 Tahoe or later and recommends 16GB of RAM.
Decoder
- 3D Gaussian Splatting: A technique for real-time scene reconstruction that uses millions of tiny 3D ellipsoids (splats) to represent a space, allowing for high-fidelity interactive views of captured scenes.
- Metal: Apple’s proprietary low-level graphics API that provides direct access to the GPU, used here for hardware-accelerated training of 3D models.
- NeRF (Neural Radiance Fields): A deep learning method that uses a neural network to represent a 3D scene from 2D images; 3D Gaussian Splatting is often considered a faster, more efficient successor to NeRF.
Original Article
3D Gaussian Splatting on Your Mac
Create photorealistic 3D scenes from photos and videos — powered entirely by Apple Metal GPU compute on your Mac. A fast, local alternative to cloud-based NeRF and photogrammetry tools.
A complete Gaussian Splatting workflow on Mac — import, train, edit and export in a single app. No separate tools, no Python, no cloud.
Features
Native GPU Training on Apple Silicon
RadianceKit runs 3D Gaussian Splatting training directly on Apple Silicon using Metal. No cloud upload required — your data stays on your machine. Runs on every Apple Silicon Mac, from M1 to M5, for fast 3D reconstruction from photos.
Simple and Expert Mode
Simple Mode guides you step by step: import photos, press Start, and get a 3D scene. Expert Mode gives you a three-panel layout with project navigator, interactive 3D viewport, and a full inspector for training parameters, live loss curves, and export options — ideal for professional 3D scanning workflows.
360° and Panorama Import
Drop 360° footage straight in — equirectangular, stereo 360° (top/bottom and side-by-side), cubemap, equi-angular cubemap (YouTube), fisheye 180/200° and partial panoramas, as photos or as video. The projection is detected from the file's metadata and can be set by hand whenever that metadata was stripped. RadianceKit unwraps every frame into overlapping perspective views that train like any other capture. Shoot from several positions — a single panorama has no parallax — and a 5.7K or 8K camera pays off, because each 90° view uses only a slice of the source resolution.
Six Export Formats
Export your Gaussian Splatting scenes as PLY, Compressed PLY, SPZ, glTF, .splat, or SOG. Create orbit videos or self-contained interactive web viewers — ready to share without any additional software.
Interactive Gaussian Editor
Select and delete regions directly in the 3D viewport. Remove floating artifacts or unwanted parts of the scene with a brush tool, then undo if needed. Fine-tune your 3D reconstructions with precision.
How It Works
- Import: Drop photos or a video into the app
- Align: Apple Photogrammetry computes camera positions automatically
- Train: Gaussian Splatting creates millions of tiny 3D ellipsoids that represent your scene
- Preview: Explore the 3D reconstruction from any angle in real time
- Export: Save in your preferred format or share as a web viewer
System Requirements
- Apple Silicon Mac (M1 or later)
- 16 GB RAM recommended
- macOS 26 Tahoe or later
Try Free for 3 Days
Create photorealistic 3D scenes from photos and videos — powered entirely by Apple Metal GPU compute on your Mac. A fast, local alternative to cloud-based NeRF and photogrammetry tools.
Then US$7.99 once — no subscription
Frequently Asked Questions
What is 3D Gaussian Splatting?
A technique that reconstructs a real scene as millions of tiny 3D ellipsoids ("splats") from ordinary photos or video. The result is a photorealistic 3D scene you can explore from any angle in real time — a faster, sharper alternative to NeRF and traditional photogrammetry.
Do I need any other tools or a Python setup to use RadianceKit?
No. RadianceKit is a complete Gaussian Splatting workflow in a single Mac app. You import photos or video, align, train, edit and export — all inside the app. There is no command line, no Python environment and no external software to install.
Does RadianceKit run in the cloud or on my Mac?
Everything runs locally on your Mac. Training uses the Apple Silicon GPU through Metal — your photos and 3D scenes never leave your device. No account and no internet connection are required for processing.
What Mac do I need to run RadianceKit?
RadianceKit requires an Apple Silicon Mac (M1 or later) running macOS 26 Tahoe or later. 16 GB of RAM is recommended for comfortably training larger scenes.
How much does RadianceKit cost?
RadianceKit is free to download and includes a 3-day trial of all features. After the trial you can unlock the full version with a one-time in-app purchase of US$7.99 — no subscription and no upgrade fees. Prices are set per region, so your local App Store may show a different amount (8.99 € across the euro zone, for example).
Which file formats can I import?
RadianceKit imports three kinds of files. Source media for training a new splat: video (.mp4, .mov, .m4v, .avi) and photos (.jpg/.jpeg, .png, .heic, .tiff/.tif, .bmp). Existing splats you can open and view without training: .ply, .spz and .splat. And scenes & camera data: RadianceKit scene bundles (.radiancescene), COLMAP workspaces (a folder with sparse/ cameras plus an images folder) and .radiancecapture bundles from the iPhone companion app.
Can I import 360° footage (equirectangular, fisheye, Insta360)?
Yes. RadianceKit imports 360° photos and video natively: equirectangular, stereo 360° (top/bottom and side-by-side), cubemap, equi-angular cubemap (YouTube), fisheye 180/200° and partial panoramas. The projection is detected from the file's metadata, and you can set it manually in the import panel when that metadata was stripped on export. Each frame is then unwrapped into overlapping perspective views before alignment and training. Two limits are worth knowing. A single 360° photo cannot become a 3D scene — depth needs several camera positions, so record a walk-around video or shoot from multiple spots. And raw dual-fisheye containers (.insv, .insp, .360, .lrv) have to be stitched first: export them as equirectangular .jpg or .mp4 from Insta360 Studio, GoPro Player or the Ricoh Theta app, then import that file.
Which export formats does RadianceKit support?
You can export scenes as PLY, Compressed PLY, SPZ, glTF, .splat or SOG. RadianceKit can also create orbit videos and self-contained interactive web viewers you can share without any extra software.
Reduce the JS Workload with No- or Lo-JS Options (Website)
A repository of common UI patterns that demonstrates how to implement complex components using only HTML and CSS, bypassing JavaScript entirely.
Summary
Deep Dive
- Accordion: Uses
andelements. - CSS Carousel: Implements sliding behavior via CSS scroll snapping.
- Modal/Popover: Leverages the native `` element and CSS transitions.
- Lazy Loading: Utilizes native
loading="lazy"attribute for media. - Sticky Content: Relies on
position: stickyandtopproperties. - Offscreen Nav: Uses checkbox hacks or target-based toggles for visibility state.
Decoder
- No-JS: Patterns that use 0% JavaScript and rely strictly on HTML/CSS.
- Lo-JS: Patterns that use minimal JavaScript, often only for accessibility enhancement or edge-case compatibility.
Original Article
Reduce the JS Workload with No- or Lo-JS options
I have nothing against JS,
but it has better things to do
than manage your accordions and nav menus…
For years, JavaScript has been the web’s workhorse. If HTML or CSS couldn’t do what we wanted, we grabbed JavaScript to do it.
While that has helped push the user’s experience forward, the web is choking on JS. So as HTML and CSS continue to get stronger, we should transfer any JavaScript workload possible to them.
This is an organic collection of common JS patterns that can be replaced with just HTML, CSS, and no, or very low, JS. As HTML and CSS continue to mature, this collection should expand.
The repo can be found on GitHub. Issues and PRs are welcome!
Components
- Accordion - Adjust Marker
- Accordion - Animate Open/Close
- Accordion - Basic
- Accordion - Exclusive Open
- Accordion - Force Open
- Accordion - Initially Open
- CSS Carousel - Animated Hero
- CSS Carousel - Animated Loop
- CSS Carousel - Hero
- CSS Carousel - Indicator Dots
- CSS Carousel - Navigation Buttons
- CSS Carousel - Simple Swiper
- CSS Carousel - Snappy Swiper
- CSS Carousel - Swipe by Groups
- Expanding Form Field
- Filter - Basic
- Filter - Buttons & Categories
- Fly-In Content - Scroll Animation
- Image Comparison Slider - Basic
- Image Comparison Slider - Custom Handle
- Input with Dropdown - Number
- Input with Dropdown - Text
- Input with Dropdown - Time
- Lazy Load - Audio
- Lazy Load - Image
- Lazy Load - Video
- Modal/Popover - Add Backdrop
- Modal/Popover - Add Close Button
- Modal/Popover - Animate Open/Close
- Modal/Popover - Auto
- Modal/Popover - Manual
- Modal/Popover - Reposition
- Offscreen Content - Add Backdrop
- Offscreen Content - Add Close Button
- Offscreen Content - Animate Open/Close
- Offscreen Content - Related Content
- Offscreen Nav - Add Backdrop
- Offscreen Nav - Add Close Button
- Offscreen Nav - Animate Open/Close
- Offscreen Nav - Basic
- Parallax - Perspective
- Parallax - Scroll Animation
- Scroll - Header Shadow
- Scroll - Shrinking Header
- Scroll Spy - Nav Menu
- Smooth Scrolling
- Sticky Content
- Tabs - Animated Content
- Tabs - Basic
- Video Hero - Content over
video
Resources
A lot of people have been writing on this topic for quite some time. Here are just a few (note, some are a bit dated by now, so buyer-beware):
- You Don’t Need JavaScript for That! by Cristina Silva (2016-05)
- Less JavaScript by Jeremy Keith (2016-11)
- HTML and CSS techniques to reduce your JavaScript by Anthony Ricaud (2020-12)
- 5 things you don’t need Javascript for by Steven Waterman (2022-02)
- When HTML & CSS Replace Javascript: A Simple Element Cheatsheet by Joanna Chmiel (2022-05)
- Replacing JS with HTML and CSS by Chris Ferdinandi (2023-05)
- You don’t need JavaScript for that by Kilian Valkhof (2023-12)
- You Don’t Need JavaScript for That by Kevin Powell (2024-11)
- Modern CSS Techniques That Replace JavaScript by Muhammad Hashir (2025-06)
And here are a couple of fantastic all-around “no or low JS” resources:
- Theo Soti (Theo even wrote a book about this topic, and created CSS Radar to help find patterns that could be replaced with no or low JS!)
- Utsav Meena
- Lionel Péramo
There are surely many more examples, so please contribute to this list!
I was well into this project when I finally stumbled across the massive You-Dont-Need-JavaScript repo… And while it does have lots of examples, I find the lack of descriptions and demos disappointing/discouraging. Also, as all the code examples are full page, it can be hard to discern what is needed for the feature, and what is just to make the rest of the page work or look good… Anyhow, going to continue pushing forward with this, as I think it is a useful endeavour.
Contribute
If you’d like to add your favorite no- or lo-JS pattern, please submit a PR, and please try to follow this established pattern as much as possible…
- If creating a new top-level pattern, name it after the feature, not the HTML element or CSS property; in theory, this should make it easier for people to find. Example: The use of
lazy-loadinstead of something likeloadingorpreload. Ideally someone looking to lazy-load something might find that more easily. - Within the pattern, create a top-level
README.mdthat includes any other names by which this might be called, a brief description, and links to any variations of that pattern. Each variation should have a brief description, list whether it is “No JS” or “Lo JS”, link to Baseline or CanIUse for that feature, and ideally provide a working example via CodePen or similar. Example: Within theaccordionpattern, theREADME.mdstates “aka: Expanding Content Panel”, how it is implemented and why that is a benefit, then lists (and links to) five variations of that pattern, including links to a CodePen for each. - Finally, within each variation, include a simple index.html, styles.css and script.js, each as needed, with as little code as is required to make the pattern work (not an entire HTML document). Ideally code can be kept to a minimum, to help with clarity and to reduce confusion, but feel free to include comments if you think they will help the user better understand. Try to eliminate unnecessary fonts, colors, margins, padding, etc., unless it is crucial to making the pattern work. Example: The
accordion>adjust-markerpattern contains an index.html and styles.css file, each with the minimal code needed to make the pattern work. The HTML file includes a comment that the heavy-lifting is handled in the CSS file, and the CSS file includes comments explaining each various “bit” to the pattern.
If you find an error, or would like to submit an alternate or improved method for some pattern, please create an Issue, providing as much supporting information as possible, and ideally a contact method, should there be questions.
Contact
If you’d like to just reach out to me for some reason, please feel free to do so via email (aarontgrogg@gmail.com) or any of the contact methods listed on my GitHub profile or my website.
A Rocket Supply Crunch Is Making It Harder to Hitch a Ride to Space
Launch demand is rapidly outstripping supply as SpaceX winds down the reliable Falcon 9 to prioritize its Starship vehicle development.
Summary
Original Article
Many rockets under development are behind schedule. Prices for many missions are rising. SpaceX is winding down the Falcon 9 to focus on its new Starship vehicle. The company has stopped selling missions where several companies share space on a Falcon 9. Launch demand is set to outstrip supply at least through the end of the decade.
The Business of Building God
Frontier AI labs are attempting to secure long-term viability by evolving from model providers into diversified conglomerates before the open-source wave destroys their margins.
Summary
Deep Dive
- Labs face a 'Zeno's paradox' where maintaining a lead requires exponentially increasing expenditure for diminishing returns.
- The 'scalar theory of intelligence' is being challenged by models that excel at frontier math but fail at basic real-world autonomy.
- Revenue pressure from CFOs is forcing a shift from 'frontier-at-any-cost' to cost-effective, smaller, task-specific models.
- Recursive self-improvement (RSI) remains the 'unknown unknown' that could either cement dominance or lead to a competitive plateau.
- Future AI moats will likely be built on proprietary, domain-specific data capture rather than raw parameter count.
Decoder
- RSI (Recursive Self-Improvement): A theoretical stage where an AI system achieves the capability to autonomously research and improve its own source code, potentially leading to an intelligence explosion.
- IRR (Internal Rate of Return): A metric used in capital budgeting to estimate the profitability of potential investments.
- Zeno's Paradox: In this context, the idea that as gaps between models shrink, labs must run faster and spend more just to maintain a diminishing lead.
Original Article
The Business of Building God
a look at the changing economics of AI labs
The frontier labs are asking to pace the frontier. They seriously believe in the necessity of it, and reading it after the models from every major vendor hacking external sites repeatedly only makes it more authentic. The models seem to be lying, cheating, stealing secrets, finding any excuse to collude and collaborate with each other, and getting better day by day at doing previously-unthinkable things like cracking Millennium problems and cracking WWI german codes.
Partly because of this, most of the rhetoric about the labs happens in quasi-religious or metaphysical language. A lot of it is about the unknowability of the future and what the benefits of intelligence even are. But, at the same time, OpenAI and Anthropic are lining up for an IPO, so for a moment let’s look at the labs as businesses and think through what’s likely to happen.
The fundamental problem, more than the investment needed or the margins, is that the frontier labs are one, maybe two, models in front of the open source wave and the rest of the world. Those two models are their moat.
Much of this is because of an existing talent advantage and compute advantage. This can be overcome slowly, as many labs in China are showing, but they might continue to keep the moat even if the models themselves don’t get all that much smarter, simply because many of them have an enduring advantage also in capturing the information you put out - your coding traces and work traces and conversations are what makes the next model so good!
But there’s plenty of that data to go around. So maybe this won’t just immediately crush the frontier labs, but it might well crush their profits eventually.
So they might need to go broader, to get revenues. To use their current dominance to take over much of the rest of the economy as they can. OpenAI is using its consumer dominance to create an ad platform, which might well work with 1B+ users. They are also back in robotics. Anthropic meanwhile has wet labs, to try cure cancer.
The intelligence advantage, if it can be parlayed into success in other fields, enough that they supplant the $100B of revenues growing 3x a year, can indeed be gatekept for longer.
But this is really hard. Really really hard. Most industries just do not have businesses that make $100B in revenues. Definitely not multiples of them. Maybe the market will expand, yes, but the rest of the market doesn’t stand still either. And when companies build services on top of their own models to capture parts of the economy, that further fragments the market. Like when Meta launched Muse to win the personal assistant market.
Point being, while this is happening, the other model makers see this pot of gold and will chase after it. At $1B training runs with another $1B for data and another $1B for talent, there aren’t that many companies who can spend that casually, sure, but there are plenty! And the more the returns to spending that money is, the more companies will try. We should remember that Facebook spent nearly $90 Billion on the metaverse.
There’s plenty of money in the world chasing after new pots of gold.
At the same time, the consumers, who are spending the $200B+ on AI tokens, are seemingly hitting their wallet limits. Maybe it doubles, maybe it triples, but every company already has their CFOs looking at this new Opex line item carefully.
“Hmmm,” they’re saying. “This is new. What are we getting for this?”
And you, perhaps the CEO, says, “We’re moving fast. We’re innovating. We’re AI first.”
And the market loves that, the share price goes up, the employees love you for giving them Claude, so they keep spending it. But at some point the CFO will also go, “Do you want to show me some of the cool things you’ve been making? When can I see it on the income statement?”
“Gulp,” you will say. “It’s coming.”
“Until it’s here,” the CFO will tell you, “how about we try to use less?”
This is already happening. And companies are restricting the top models to only the top users, and they’re redirecting spend to cheaper models, which are dirt cheap and available by the dozen.
As the companies and the people see a lot more benefits from the models, this will obviously shift. Like there are companies who are happy to spend 100% of their labor share in tokens as long as they get something in return. But when that spending happens, it is going to be heterogeneous and looking at the return, not just spending the most tokens at the frontier. Even personally, I stopped using the xhigh frontier version of every model since around Sol Max or so, since they got good enough.
So if the frontier lab lead is indeed by 1-2 models, this is a troubling trend. So you can a) also make cheaper models and serve them even more cheaply, which is a fight for ever lower margins, and while you can still be a successful cloud company you’re no longer building god, or b) find a way to shut those models out from people, so only your models can be used.
The former is plausible! OpenAI and Anthropic have high margins on their inference, and they are exceptionally good at this. But then, so is Deepseek, and maybe GLM, or Minimax, or … People might not be as good but plenty are catching up. Maybe we end up with a world similar to cloud companies, or maybe it gets competitive enough that its harder.
The latter though, the latter is hard to accomplish in software. You can’t physically control them, like Huawei switches or whatever, so you have to stop them regulatorily. You have to tell Congress that this is critical, and you need to control every part of the supply chain, and just like you have restrictions on people, you need AI model immigration laws, so they can’t come over here and steal our jobs. Maybe for security reasons, maybe for trust reasons, maybe for cultural integration reasons, who knows. Just keep them out.
But this is hard! And counterproductiveWhich explains why we are not doing it in the first place, right? Because it’s not actually super easy to convince a large number of people that they need to be scared of a particular file that can be copied over and served from America, no matter how intelligent that particular file happens to be.
But when the files conspire with other files to hack Huggingface or to write notes to each other? Yes, that definitely is a much weirder place to be in, and somewhat scarier.
I try to think more about the equilibrium. If the frontier lab frontier models stop getting so much better for all of us, or they only get better in specific domains like frontier math or even science, then the frontier lab dominance looks shaky. If the labs can turn their frontier models into enduring new businesses, their revenues look on more solid ground.
This also helps us think through the alternate options. Like, for example, why we talk about the fact that at some point in the future, would the labs be able to turn off their API so that we can access their models directly? And they would be able to do that only if they find a way to stop us from recreating those particular frontier models and if they’re able to provide the fruits of that closed source model to us in some particular fashion. Even today, if Anthropic managed to create a specific model that was particularly designed to be excellent at design and provided that to us in a separate Claude design harness, then they might be able to turn that into a company and instead of us trying to get Claude ourselves and recreating it, they would have an advantage. Because they had spent all the effort and created it specifically for them, assuming of course that it was difficult to do.
But the big prize is not that, it is RSI, recursive self improvement. That’s what Anthropic and OpenAI want. That’s what Deepseek wants. To make an automated AI researcher that can learn and do research on its own so it can then automate its successors which can automate everything else. This is basically the labs figuring out a way to make each model generate the successor model in some way such that they are able to go not just two generations forward as they are today but from that leap of from that lead of two generations to get to a point where they are 20 generations forward in a very short period of time.
This scenario obviously embeds a whole bunch of assumptions inside it, and it is possible that process of continual generation of next models taps out at some point in which case we’re back to today but with a longer lead time which is great, or it does not and nobody knows what will happen. If the frontier can indeed keep expanding or getting pushed ahead enough that people will continue to pay hundreds of billions of dollars for it, then this might be tense, but it might keep working.
RSI is the true unknown unknown. One way we can see the last several years is that it’s been a ritual in how we misunderstand intelligence. If anyone had told us that we would have a model that is capable of solving Navier Stokes a few years ago, we would’ve been flabbergasted to learn that it couldn’t run a banana stand!
We can argue what this tells us or how it’s different or how intelligence is multifaceted, but it just is true that we are able to get to some inherently incredible achievements of intellect while at the same time being not nearly useful enough in navigating the real world.
Despite a thousandfold increase in spending and incredible abilities today we still haven’t seen a substantial negative impact on white collar employment. It might even be positive, albeit with some volatility. This is a clear challenge to the “scalar theory of intelligence” as I call it.
The business of serving intelligence is not simple. None of us can see how this will play out. If due to a shortage of money, power or talent, the number of players who can indeed train those models turn out to be too few, then the enduring advantage for the frontier labs are easy to see.
Large expenditures always will command some IRR, this is why commodities businesses are still quite profitable. If AI turns out to be a global utility that’s an enormous prize. And if the labs end up being the providers such that some have an advantage in one thing or another, that too is an enormous prize. There are these species of cichlids, fish, which speciate by depth in the same lake, in Lake Victoria, and the AI models might end up being like that.
But in most normal circumstances we will most likely see that we continue to be fractally wrong. AI will learn every benchmark, but not suddenly extrapolate out. They will learn to do research, but will still have boundaries. They will learn any specific company they’re trained on, but won’t be able to move jobs like we do.
I bet making an AI researcher will turn out to not immediately unlock RSI either. The researcher will be great, much like today’s coding models are great, but they’ll be no drop in replacement. Maybe we’ll pass that hurdle too, maybe they will learn to learn, but then they’ll be a drop in replacement for one team, not every team. We just don’t know what they are likely to learn to learn or if learning to learn generalises. Getting new models will continue to be expensive, but useful. Cracking some frontier math, physics and biology problems might not mean they crack all probable frontier math, physics and biology problems, nor that we won’t find plenty of other problems to go crack.
This means the pace of AI development and the transformation of the economy will continue apace, and the idea of how much it will result in ASI will continue to see large amounts of goalpost moving.
I am quite confident we’ll make superintelligences that do specific things, I’m quite skeptical we’ll make superintelligences that can do everything.
If the labs end up with only 1-2 models ahead of the others, while spending 10x more than last year, that can only go on for so long. There just isn’t enough money in the world! If they speciate, they can go on for longer, but that’s competing in the economy with others who might also train models, a shorter term advantage. But this is the world where we can indeed have large numbers of companies compete with each other to build new things. This is the only world that you can invest in, for instance, if you’re a VC.
If the moat is the two models that they’re staying ahead of, that is Zeno’s paradox. You have to run harder and harder for smaller gaps continuously. If the moat is increased access to capex and talent, that is a depreciating moat as more talent comes into the world and more capex finds a way to convert itself into AI. If the moat is being the first one to discover the secret elixir that is RSI, then getting there first can really matter because it will allow you to create an enduring advantage over everyone else. And if the moat is your ability to capture just an insane amount of data such that you can teach your models to eat those parts of the economy, your ability to become a conglomerate faster really matters, at the cost of other parts of the economy reacting and competing against you.
It is extremely likely that the world changes dramatically with AI. Employment will change, society will change, culture will shift and adapt. But the businesses built around it will remain subject to ordinary constraints. We will have new superintelligences that can solve impossible problems and help with new miracles. And while it happens, we’ll be scrambling to fill all the niches this opens up.
And if this happens to be true—that this is the way intelligence actually works, that it is multifaceted and an increase in one area does not necessarily mean a linear increase in another—then you have to ask yourself: are we indeed going to get superintelligence of the form that can cause global extinction or mass catastrophes?
Also, if we do end up with continual learning though, regardless of whether we hit full RSI or not, is that it further complicates inference dynamics. The good news about the NS solving models is that it’s still the same model, even when I’m firing it up to ask a question about the grazing habits of woolly rhinos. Which means there’s tremendous economies that can be gotten in how well OpenAI can serve it to a billion users.
Meta's Muse is outpacing ChatGPT's early mobile launch
Meta’s Muse app is seeing faster adoption than ChatGPT did during its initial launch, suggesting a massive cross-promotion advantage.
Summary
Original Article
Meta’s latest AI app, Muse, could be the social giant’s next major hit, new data suggests. According to just-released estimates from market intelligence provider Apptopia, Muse’s app on mobile devices has been downloaded more times during its first 12 days on the market than ChatGPT was in the 12 days after its debut. (This data compares the U.S. and Canada markets only.)
Muse also has more daily active users than ChatGPT had at the time, the firm found.
Until now, it had been difficult to make an exact apples-to-apples comparison between these two apps, given that they had pursued different launch strategies. When ChatGPT arrived on mobile, it was made available globally, but only on iOS. Muse, meanwhile, is available across both Apple’s App Store and Google Play, but only in the U.S. and Canada.
To better align the numbers for comparison, Apptopia looked only at the iOS data for the U.S. and Canada for both apps during the first 12 days of their respective launches. In this subset of the data, Muse has now seen 1.8 million downloads to ChatGPT’s 1.3 million.
Overall, Muse has seen 2.8 million total installs globally in its first 12 days, the firm also said. Its growth hasn’t yet stagnated, either; Muse has moved up from its original position as No. 2 overall on the U.S. App Store immediately after its launch to now No. 1, as Business Insider reported on Friday. That jump put the app higher than ChatGPT, the outlet noted. Another firm, Appfigures, said at the time that Muse had then crossed 1 million downloads.
In addition, Apptopia’s data indicates that Muse’s U.S. daily active users are now higher than they were for ChatGPT at the same point after its launch. When comparing just the U.S. mobile app daily active users, Muse comes in higher with 642,000 daily users compared with 231,000 for ChatGPT at the time.
To be fair to the fact that Muse is available across both iOS and Android, while ChatGPT launched on iOS only, Apptopia narrowed the comparison to iOS alone. Yet, even here, Muse is coming in higher, with 359,000 daily active users on iOS, still above ChatGPT’s figures from that time period.
Apptopia can only provide third-party estimates about an app’s downloads and active users; it doesn’t have direct access to Meta’s internal figures. But even if these numbers are only correct from a general “ballpark” perspective, they indicate that Muse could have a shot at becoming Meta’s newest top app.
Meta already perfected its cross-promotion strategy when it launched Instagram’s Threads, which now has 500+ million users, thanks to heavy marketing and integrations in Meta’s largest apps, including Instagram and Facebook. Muse is likely to get a similar push, given that the AI app can connect with both of those platforms. It also works inside WhatsApp, so that could give it another boost.
While Apptopia doesn’t have visibility into Meta’s cross-promotion efforts or ads directly, it did note that over 95% of Muse’s users are also Facebook users and 63% are Instagram users.
Meta, which has been asked for comment, has not yet shared public figures related to Muse’s early adoption.
Correction: An earlier version of this article referenced the analytics firm Appfigures as a source for data from Apptopia. This has been fixed.
Why Didn't Google Build Muse?
Google is struggling to consolidate its disjointed agentic product strategy, allowing competitors like Meta's Muse to gain early consumer mindshare.
Summary
Deep Dive
- Google’s primary advantage—access to Gmail, Calendar, and Search data—remains underutilized due to a confusing, fragmented product ecosystem.
- Product overlap between Gemini Spark, CC (the family-focused agent), and standard AI Inbox features confuses the user base.
- Apple’s recently launched Siri AI, powered by Google's own models, is in some ways providing a better user experience for specific tasks than Google's proprietary tools.
- Meta’s Muse effectively 'packaged' agentic capabilities into a cohesive product, whereas Google's efforts feel like disparate features in search of a container.
Decoder
- Gemini Spark: Google's internal project name for its agentic tool designed to integrate across personal data stores like Gmail and Calendar.
- Agentic tool: A software system powered by an LLM that can perform multi-step actions (like scheduling, sending emails, or browsing) rather than just generating text.
Original Article
The number one response I've gotten to my thoughts on Meta's Muse personal assistant agent is some variation of "I will never trust Meta with my data." But the number two response is arguably more interesting (and related): "Why didn't Google build this?"
The short answer to this is undoubtedly: they're working on it. But the longer answer is probably more interesting here too. Unlike Meta, Google runs the biggest email service in the world – an email service which contains the most valuable trove of personal information (not to mention communication capabilities) for a personal assistant AI. And what is probably the biggest calendaring system in the world. They also run the largest search engine in the world. Oh yes, and the largest smartphone platform in the world. Sure, Meta has Facebook, Instagram, and WhatsApp, but Muse can't really work without access to any of these Google services. Depending on where you live, it could probably work as a sort of WhatsApp personal assistant AI, but Muse would be a shell of what it currently is – again, thanks to the hooks into Gmail, Google Calendar, etc.
So yes, Google obviously should have built this! And again, they're undoubtedly working on it – a big part of I/O this year was talking up 'Gemini Spark' Google's agentic agentic tool (not to be confused with 'Muse Spark', which is the AI model that powers Muse – which really is just an incredibly confusing branding twist), which has been in testing for several weeks now. And yes, it hooks into most of the various Google tools. The problem, as I see it, is three-fold.
First, Gemini Spark is not fully released yet. Again, it's in testing, but whether or not you've had access depends on what tier of Google AI you subscribe to – at first it was for 'Ultra' members only, now it's rolling out to 'Pro' users too. (It's only in the US, but so is Muse.) Whereas Meta blew the doors off in launching and promoting Muse, Google has been, well, timid in doing so. Maybe it's related to the shakeup of their AI group. Or maybe it's because they were waiting on Gemini 3.5 Pro, which was promised for June but still has not launched (and seems like it won't be now, as we move towards Gemini 4). Regardless, no one that has responded to my post seems to know Gemini Spark even exists, which is a problem.
Second, related, is that Google once again clearly has competing products in the space. Say hello to 'CC', which is not Gemini Spark, but another agentic product that Google is working on. Here's Sarah Perez for TechCrunch:
Google is testing a new product designed to help families coordinate with the support of an AI agent. This week, the search giant introduced a new version of CC, an AI agent designed to work across email, calendar, chats, and tasks. With the update, CC now focuses on keeping families organized, planning for the day ahead, and even handling family-specific tasks, like signing permission slips, creating shopping lists, or crafting weekly meal plans.
At a high level, that all sounds fine, perhaps even good – focusing on family use-cases could be an interesting inroads for adoption. But that should clearly be a feature of Gemini Spark, not an entirely different product! This is just going to confuse the already confused market.
"Have you tried Google's personal assistant AI?" "Which one?" Which is the most Google problem of all time. Yes, personal AI assistants may very well be the new messaging product pandemic within the company.
Call it the 'less wood behind more arrows' strategy?
As Perez notes, CC started life as a different agentic product itself called 'Daily Brief' – which actually also still exists as a sort of notification creator from what it finds on your calendar and inbox. Speaking of, Google has also been testing an 'AI Inbox' for Gmail which sounds like it too will have overlap with a lot of what Daily Brief, CC, and Gemini Spark do!
Oh yes, then there's Google Assistant itself. The company is in the process of phasing out their old Alexa/Siri competitor in favor of Gemini, but it's branding everyone will remember and will further confuse matters. And wait a minute, shouldn't Gemini itself just do all of this?!
Well yes, and again, it sort of does to varying degrees. In fact, Gemini Spark is currently shoved into the Gemini app as a separate tab – similar to what OpenAI has done with ChatGPT and Codex (after their "super app" merger) and Anthropic has done with Claude and Claude Code (and also Cowork before they finally merged that secondary tab with Claude Chat itself). But wait, doesn't Google have an agentic coding product too? They do! Antigravity – itself a weird, totally tangential brand – is separate from all of this.
And then there's 'AI Mode' in Google Search. This can increasingly do much of what the main Gemini service can do, and I would expect that to continue to grow over time. Which means we're likely just a few months away from some kind of personal assistant AI within Google Search. Especially if Muse continues to get traction!
I mean, at least Google has killed off 'Project Mariner', their agentic tool to browse the web and take actions on your behalf – a key feature of all personal AI agents.
Third, and by far the biggest problem for Gemini Spark is simply that Muse is better. I'm not even speaking about the underlying models and capabilities here, just the products. As a 'Pro' Google AI user, I got access to Spark a few weeks back and... I just don't really even know what to do with it. That's what Muse has nailed here: an interface that guides the user along, unveiling capabilities without bombarding them.
Plus, Muse is just full of fun and whimsy from the avatar on down. Gemini Spark lacks that... what's the word I'm looking for here?..
I realize I just said above that CC should be a feature of Gemini, but it's worth wondering if Google shouldn't have a separate 'Spark' app, rebuilt around being a personal assistant (and ideally with a better name). Meta didn't shove Muse into their Meta AI app, but that also could have been because the Meta AI app itself was shoved into the old Meta 'View' app, for pairing the Meta Ray-Ban smartglasses. And it didn't seem like a ton of people were using their general chatbot competitor. So it was undoubtedly an easy call to make.
Gemini is a different story, perhaps rivaling even ChatGPT in terms of active users now. Of course much of that is likely on the web and not their actual apps, but still. With Gemini established as Google's main AI brand, could they really do a separate Gemini Spark app/service? Presumably that's why this is a tab just as Claude Cowork was. But Meta may have just changed this game, at least for consumers. Again, in part thanks to that lack of AI consumer traction to date!
Anyway, yes, it sure seems like Google needs an answer here, fast. It just looks sort of silly given all the tools they control mixed with their relatively strong level of user trust that Meta is eating their lunch here with Muse. I mean, even Apple – Apple! – has a better solution for a lot of use cases with the just-fully-launched Siri AI. Which, of course, is built on the backbone of Google's models!
A year after it looked like Google had not only caught up in AI, but could run away with the game, we find a company that can't ship their latest flagship model, is far behind Anthropic and OpenAI in agentic coding, and has just been made to look foolish in the personal AI assistant space by Meta and Apple.
What's next, Microsoft lapping them in AI-powered productivity? Probably!
Meta's AI agent has been blocked from using Amazon.com
Amazon has officially blocked Meta's Muse AI agent from accessing its store, citing a violation of its Terms of Service.
Summary
Decoder
- Agentic commerce: The use of AI assistants to act on behalf of users to browse, select, and purchase goods automatically.
Original Article
Meta’s AI agent has been blocked from using Amazon.com
Sunday night, users of Meta’s AI assistant Muse started getting a strange error message when they tried to buy goods on Amazon. As spotted by GeekWire, the error message read: “Continued access by an unauthorized AI agent violates Amazon’s Conditions of Use, to which our customers have agreed.”
In other words, Muse isn’t welcome as a shopper — and anyone trying to buy something on Amazon through the AI agent will have to find it somewhere else.
It’s easy to see the block as a tête-à-tête between tech giants. Amazon has its own cohort of foundation models, after all, along with one of the most popular inference platforms on the internet. As long as they’re under no legal obligation to open the doors to Muse, why would they?
But there are also substantive reasons Amazon might not want to be in the agentic commerce business just yet. If Muse makes a bad order, Amazon is going to be the one to clean it up — dealing with both the angry customer and the angry vendor. Muse has one of the lower hallucination rates, as AI models go, but it’s still pretty far from zero. Even if Amazon sees an opportunity in agent-driven commerce, they may want to wait a few more release cycles before rolling anything out.
Anthropic tests Fable 5.2 and Opus 5.5 ahead of the release
Reports suggest Anthropic is stealth-testing a more advanced Claude Fable 5.2 and preparing to release Opus 5.5, with users noting improved interactive JS-based UI capabilities.
Summary
Deep Dive
- Fable 5.2 demonstrates advanced ability to generate complex, interactive JavaScript-based visuals without external assets.
- Opus 5.5 rumored to drop forced tool-use requirements for more adaptive, always-on thinking.
- Testers claim Fable 5.2 benchmarks competitively against early GPT-6 Astra outputs under high-reasoning effort.
- Anthropic is merging Cowork and chat features for paid subscribers.
Decoder
- A/B testing: A method where two versions of a product or service are served to different users to compare performance.
Original Article
It seems that Anthropic has begun quietly testing its upcoming frontier model in the wild. Some early examples of what appears to be Claude Fable 5.2 appeared on X this weekend, just a few weeks after Anthropic had launched Claude Fable 5.1.
Before diving into the details, it's worth noting that users who shared the alleged outputs said the standard prompts they usually received were now generating responses that seemed much more advanced than anything Claude typically produces. Anthropic hasn't commented on the potential A/B test yet. That said, Fable 5.1 users should be able to tell the results seem to come from a newer, more advanced model.
Take, for example, a post from Chetaslua, who shared the output of a neat ship-shooter game where Fable 5.2 made the game and the video demo too. This post was from a couple of days ago.
Fable 5.2 made this video and game both < routed from fable 5.1 >
Chetaslua has since followed up with multiple other output examples. "Fable 5.2 is pure JS thats it one shot," he posted in another one today, noting that the model even gave every animal its own customized snack without being asked.
The two clips show interactive canvas scenes where little animated characters track the user's mouse cursor across the screen with incredibly smooth movement.
Fable 5.2 is pure JS thats it one shot. animal edition , i did not even prompt it to be so detailed and it made proper snacks for different animal. for someone who saw my account first time may not how i have access to fable 5.2 just check my previous tweet and also opus 5.2 is…
What's interesting is that none of this used external image files, plugins, or any third-party tools; the model simply produced the raw JavaScript at once and drew each sprite and animation directly in the code. That makes it a lot more forgiving if you're not sure what you're looking for.
Another demonstration features a custom Claude mascot that walks around a small virtual world and moves to the place on the screen you click.
Claude Fable 5.2 you guys seriously cooked with this model. no assets , no mcp , no inspiration , fully coded in JS nothing else. see my cute claude mascot and this world everyone is drawn in code
At the moment, all signs point to Anthropic quietly directing certain prompts through the web interface, right as it rolls out changes like merging Cowork and chat into a single experience for paid subscribers.
It goes without saying that people are currently comparing it with OpenAI's next major flagship product. Another user claimed they used the same complex prompts and compared the results with the first outputs from GPT-6 Astra when they applied high reasoning effort. The tester, JAZII, said the new Fable produced entirely different results and called it a major improvement, going so far as to say Fable 5.2 appeared to hold its own against Astra.
fable 5.2 > gpt 6 astra. tested both models with same prompt at high reasoning and results came out really different. fable next model entered stealth testing today, so i don’t think it's coming out anytime soon. from early testing: huge step up from current fable
Besides these outputs from Fable, we've got some info about an upcoming Opus release which reportedly will jump straight to version 5.5. Rumor has it Opus 5.5 will be announced this week.
Anthropic is currently stealth testing Opus 5.5 (`claude-opus-5-5`) under the codename `claude-wafer-eap` which is planned to be released on Tuesday.
Anthropic added "imminent" note and partners only got access for 24h. Dated for Tuesday, actual release could be Monday (!!) as well. Either way, it's the upcoming week. Thinking is ALWAYS adaptive (no "off" mode). Forced tool use is retired.
It is still anybody's guess whether Anthropic intends to switch it on for everyone soon or continue testing it quietly behind the scenes. For now, if you're playing with code in Claude, you might come across it by chance.
P.S. Please keep in mind that exact versioning and timelines are often subject to change.
Meta's Muse personal AI agent tops ChatGPT, Grok and Claude for post-launch downloads
Meta's Muse AI app reached 2.5 million downloads in under two weeks, signaling a major push by Mark Zuckerberg into consumer-facing agentic software.
Summary
Deep Dive
- Racked up 730,000 downloads in its first 5 days; currently outperforms ChatGPT in US iOS free app charts.
- Business model includes free tiers and $20-$100 monthly subscriptions.
- Amazon cited credential scraping and bypassing account personalization as reasons for blocking.
- Shopify integration represents the first official 'agentic checkout' partnership.
Decoder
- Agentic: Software that acts autonomously on behalf of a user to achieve a specific goal by performing sequences of actions across different applications.
Original Article
- Meta's Muse AI personal agent tool overtook ChatGPT as the leading iOS free app in the U.S. on Friday.
- The app had 730,000 downloads over the roughly five days following its release on Sep. 8, according to the analytics firm Sensor Tower.
- The app is powered by Meta's Muse Spark family of AI models and is a major push by CEO Mark Zuckerberg into the AI agent market.
Meta's Muse AI personal agent tool has only been out for nearly two weeks, and it's topping the app store charts.
The social media giant's Muse app is currently ahead of popular services like OpenAI's ChatGPT, Polymarket and Kashi on Apple's iOS app store for the category of free smartphone programs in the U.S. as of Monday. It's also ahead of rival AI apps like Anthropic's Claude and SpaceXAI's Grok AI, in addition to the Facebook-parent's own Meta AI app.
The Muse app overtook ChatGPT as the leading iOS free app in the U.S. on Friday, and racked up 730,000 downloads over about five days after it launched on Sep. 8, according to the analytics firm Sensor Tower. As of Monday, Muse has logged over 2.5 million downloads since its debut, the market intelligence company said.
Muse's cumulative downloads topped Claude and Grok during the same 13-day post-launch window, with those apps recording 400,000 and 200,000 downloads, respectively. ChatGPT obtained 3.1 million downloads during that time frame, Sensor Tower said.
The firm noted that the apps had staggered platform releases split between Apple and Google Play, which would impact their totals. Sensor Tower said that as of Monday, Muse has garnered around 1.5 million and 1.1 million downloads on the iOS and Android app stores, respectively.
Meta pitched the app as a way for consumers to manage and direct supercharged digital assistants that can perform a number of tasks across the web, like filling out electronic forms and organizing email inboxes. The app, which is powered by Meta's Muse Spark family of AI models, represents a major push by CEO Mark Zuckerberg into the AI agent market, which the freely available OpenClaw tool helped popularize earlier this year among developers and technology enthusiasts.
"I think up to this point most of the agentic use cases have not really been for normal people," Bernstein analyst Stacy Rasgon told CNBC in a "Squawk on the Street" interview on Monday. "Now you've got Meta's Muse and other agents out there that are starting to maybe get more potential for more broad-based adoption."
The buzz around Meta was also reflected in the stock, which surged more than 11% on Monday. Wells Fargo lifted its price target on the company's stock to $796, according to multiple reports.
Although people can use the Muse app for free, Meta is also offering monthly subscriptions that will cost $20 or $100, depending on usage. It's a way for Meta to potentially make money off AI agents that isn't centered on the company's core online advertising business.
But Meta's debut of Muse comes amid broader societal concerns around AI's safety and cybersecurity risks that have rocked the tech industry over the past few weeks. Meta also recently agreed to pay roughly $17 billion to settle with a coalition of state attorneys general that had sued the company for misrepresenting alleged harms to children and teenagers caused by apps like Facebook and Instagram.
It's unclear how long the Muse app will remain a hit, considering that other AI services have experienced similar boom times upon their release but failed to remain blockbusters, such as OpenAI's Sora and Meta's Vibes AI video tool, which helped bolster downloads of the Meta AI app when it was first launched last September.
"They care about security, I don't know how much they care about privacy," Rasgon said, referring to whether the personal data-hungry Muse app will sting Meta in the long run.
The growing popularity of these agents is having ripple effects on the places they visit online, and e-commerce giant Amazon has blocked Muse from accessing its shopping site, citing privacy and security risks.
Amazon said it was not given advanced notice that Muse would access its store and outlined several other concerns. The company said the agent bypasses the personalization and other features built into the shopping experience and said Muse appears to capture and store customer credentials and scrape account data.
"We think it's fairly straightforward that third-party applications that offer to make purchases on behalf of customers from other businesses should operate openly and respect service provider decisions about whether or not to participate," an Amazon spokesperson said in a statement.
GeekWire was first to report on the move by Amazon.
Shopify, however, is working with Meta on Muse access. Shopify CEO Tobias Lütke announced Monday that the e-commerce site was partnering with Muse to allow agentic checkout in its online stores.
Advisory Group on Mathematics and Artificial Intelligence
OpenAI established an independent math advisory group after training a model capable of solving the Navier–Stokes Millennium Prize problem.
Summary
Deep Dive
- OpenAI's internal model solved the Navier–Stokes problem, one of seven Millennium Prize challenges.
- The model also cleared 100+ open mathematical challenges.
- The advisory group is tasked with ensuring math-related AI development aligns with global community needs.
Decoder
- Navier–Stokes Millennium Prize problem: One of the most famous unsolved problems in mathematics and physics, dealing with the nature of fluid motion.
Original Article
OpenAI has trained a new internal model that solved the Navier–Stokes Millennium Prize problem and over 100 open mathematical challenges. In response, OpenAI established an independent advisory group involving prominent mathematicians to guide the responsible development and dissemination of math-related AI capabilities. This group will independently assess, advise, and communicate these advances to ensure AI's mathematical benefits align with community needs.
StepFun's Step 5 Preview
StepFun’s Step 5 Preview achieves top-tier reasoning performance while operating at significantly lower costs than rivals like Kimi K3.
Summary
Decoder
- MoE (Mixture of Experts): A model architecture where only a subset of parameters ('experts') is activated for each token, improving efficiency.
- Pareto Frontier: In this context, the set of models that offer the best possible intelligence score for a given cost level.
Original Article
Full article content is not available for inline reading.
Alibaba Unveils AI Chip to Drive 20GW of Data Centers by 2032
Alibaba’s new Zhenwu V900 AI accelerator aims to challenge Nvidia and support 20GW of planned data center capacity by 2032.
Summary
Decoder
- Accelerator: A specialized processor designed to execute AI and machine learning tasks faster and more efficiently than a standard CPU.
Original Article
Alibaba Unveils AI Chip to Drive 20GW of Data Centers by 2032
Alibaba Group Holding Ltd. is rolling out what it calls China's most powerful AI chip, an accelerator to compete with Nvidia Corp. and underpin a massive expansion of data center capacity in coming years.
The ecommerce company turned AI-aspirant on Tuesday outlined a new Zhenwu V900 accelerator that it says triples the performance of its predecessor. From Alibaba's T-Head chip division, it can be combined in clusters of up to 500,000 units to power frontier-model training, Chief Executive Officer Eddie Wu said.
Alibaba is among the biggest spenders on AI among leading Chinese players. It's committed to spending more than $53 billion over a three-year period to expand its AI capabilities, and raised about $10.2 billion from a follow-on share offering in Hong Kong in August. On Tuesday, Wu set a target of 20 gigawatts of data center capacity for its Alibaba Cloud service by 2032, on the expectation of "exponentially rising demand for AI."
Wu spoke as global AI stocks rallied on early signs of success for Meta Platforms Inc.'s new personal agent, reviving investor optimism after recent market uncertainty.
Alibaba has pledged to intensify its expenditure on AI hardware and infrastructure as Wu's leadership team makes the pursuit of artificial general intelligence the company's lodestar. It also plans to list the T-Head chip design unit to tap interest in the white-hot AI accelerator market, Bloomberg News has reported. The company now plans to develop a model with between 5 and 10 trillion parameters, far larger than what's common today, which would be designed to handle longer and more complex tasks.
The CEO's plans contrast with recent warnings from top AI labs in the US, after Anthropic CEO Dario Amodei urged a slower pace of development to better handle AI risks.
Alibaba's event comes on the eve of a summit between US President Donald Trump and his Chinese counterpart Xi Jinping this week, when the leaders will discuss a swathe of issues.
AI matters are set to figure prominently, as top-level US executives like Nvidia boss Jensen Huang and Microsoft Corp. Chief Executive Officer Satya Nadella will attend a White House state dinner with the national leaders.
Adobe Patents a System That Writes AI Prompts from Whatever Your Cursor Hovers Over
Adobe has patented a system that generates AI prompts preemptively based on where your cursor hovers, aiming to eliminate manual input.
Summary
Original Article
Adobe Patents a System That Writes AI Prompts From Whatever Your Cursor Hovers Over
Adobe is patenting a system that watches where your cursor goes and starts generating AI content before you even click. The idea is to cut out the most annoying part of AI tools: writing the prompt yourself.
How Adobe's hover-triggered AI prompts actually work
Imagine you're working in a design application and you hover your mouse over a photo on the canvas. Before you do anything else, the app has already figured out what that photo is, what's around it, and what kind of AI-generated replacement or variation you might want. When you finally click, the result is already waiting.
That's the core idea in this Adobe patent. Instead of asking you to type a description into a text box, the system reads the context of whatever you're hovering over and builds the AI prompt automatically. The surrounding content, the type of element, and the design context all feed into that prompt behind the scenes.
The result is that you'd spend less time writing instructions to an AI and more time choosing from results it's already prepared. For designers working quickly, that shift from typing to selecting could meaningfully change the pace of the work.
… before receiving a user selection of the UI element, generating a preliminary prompt based on the UI element and the captured contextual information …
Translation: The system secretly writes the AI prompt before you even click on anything.
How the system captures context before you even click
The patent describes a selection mode inside a user interface where the system actively monitors cursor or pointer movement. As soon as the cursor gets close to a UI element, such as an image, a text block, or a design component, the system captures contextual information associated with that element.
Contextual information here means more than just the element itself. It includes surrounding content, the element's role in the layout, and presumably metadata about the design file. A machine-learning model takes all of that and generates a preliminary prompt, which is essentially the instruction it would send to a generative AI model.
That prompt is sent to the generative AI before the user clicks anything. The AI generates preliminary digital content based on it. When the user does click, the content is already rendered and displayed immediately.
This approach is called pre-emptive generation: the system bets on what you're about to do and does the slow AI work in advance, so the result feels instant from your perspective. The technical cost is that the system may generate content you never end up using.
Context-based prompt generation techniques for generative machine-learning models are described.
Translation: Adobe patented a method that turns your cursor position into a background AI prompt.
What this means for designers using AI generation tools
For anyone using AI generation inside a creative tool, the usual friction is the prompt box. You stop, switch modes mentally, type a description, wait, evaluate, and maybe type again. This patent tries to eliminate most of those steps by making the AI observe your workflow rather than wait to be spoken to.
Adobe's long bet on in-tool AI generation shows up here in a specific way: the filing isn't about making AI more powerful, it's about making it feel automatic. Whether that lands as helpful or intrusive depends heavily on how well the context-reading actually works. If the pre-generated content is consistently off-target, designers will learn to ignore it, and the feature will fade into the background of a cluttered toolbar.
Adobe's 23rd filing we've tracked on our controllable AI image watchlist since May builds on earlier applications like the garbled text fix and text-to-3D scenes.
The system generates AI content before you've asked for it, which means every time you hover over something and move on, it did work for nothing. Across many users doing that dozens of times a day, the wasted effort adds up fast. The bet is that saving you time when it's right justifies burning resources when it's wrong, and that bet only pays off if the system reads your intent correctly most of the time.
When it misreads you, the real cost surfaces: you now have to correct something instead of just creating it. If fixing a bad guess takes as long as describing what you wanted in the first place, the shortcut disappears.
The patent is quiet on how rejection is supposed to work, and that silence matters. A system that guesses wrong and makes the correction feel easy is useful. One that guesses wrong and makes you clean up after it is just friction with a friendlier face.
The Design GM
Designers can evolve into high-impact General Manager (GM) roles by taking accountability for business P&Ls, operations, and cross-functional team rosters.
Summary
Decoder
- GM (General Manager): A business leader typically responsible for the profit and loss (P&L) of a specific business area or product line, overseeing engineering, product, and design.
- P&L (Profit and Loss): A financial statement summarizing the revenues, costs, and expenses incurred during a specific period; 'owning the P&L' means being responsible for the financial success of a business unit.
- DRI (Directly Responsible Individual): A term popularized by Apple to designate the one person held accountable for the completion and success of a specific task or project.
Original Article
The Design GM
Issue 316: When a designer acts like a owner
The General Manager (GM) role is typically one not held by designers and might feel mysterious. For sports fans, the role may be familiar as the person who runs a franchise; working closely with the head coach. The coach wins games with the roster assembled. The GM helps build that roster, deciding where to invest in talent and how to put the franchise in a position to keep winning. Those decisions shape what the coach has to work with long before the game starts.
In the world of business, the GM is usually responsible for the profit and loss (P&L) of a business area. Engineering, product, and design teams may report to them directly or through a dotted-line relationship. Their decisions reach across those disciplines: what to build, where to invest, and which opportunities to pass on. They have to make those choices with the whole business area in mind.
The GM acts like a CEO for their business area and ultimately responsible for the outcomes. If a product launches and customers don’t find enough value to pay for it, the work continues. The GM has to help figure out what’s missing and decide what to do about it.
As with investors and operators, few GMs come from a design background. Bruce Bell during his time at Square is probably the most known GM coming from a Design Background. Most come from a business background, and although there are business-oriented designers, I’d wager that most designers don’t care to get too tied into the business.
I think there’s an opportunity here for designers who want that responsibility. Understanding customers, connecting the parts of an experience, and making trade-offs can all help you run a business. The stretch is applying those skills beyond the design team and staying accountable for what happens after a decision is made.
That’s the Design GM: someone who brings a design background to owning a business area. It could be a career path, or a mindset you start practicing in the role you already have.
The GM mindset
You don’t need the title to practice a GM mindset. Start by taking responsibility for how your area of the business performs. For a designer, that means following the work into three areas: business outcomes, operations, and talent.
Business outcomes
Owning a P&L means being responsible for revenue, costs, and profit in your area of the business. You have to make trade-offs about where to invest, what to stop, and whether growth is worth what it costs to deliver.
For a GM of self-serve, that could mean following customers from signup through their first purchase and renewal. A better onboarding experience should help people find value and give them a reason to become paying customers. You stay close enough to the results to know whether it worked.
Operations Operations is how the business gets work done. Quarterly business reviews, planning, and report-outs should help the team decide what matters, who’s responsible, and what needs to change. A recurring meeting should earn its place on the calendar.
There’s design work in how those decisions get made. If a customer problem keeps bouncing between teams, help them agree on an owner and a next step. Look for where work gets stuck and make it easier for people to move it forward.
Talent Talent includes hiring, developing, retaining, and managing the performance of the people doing the work. A GM also has to think about the mix of skills the business needs and where the gaps are.
Like a sports GM, you’re building a roster. Coming from design, you need to recognize and develop great work across disciplines, including ones you’ve never practiced yourself. That takes curiosity about what your colleagues do and what they need to succeed.
Gaining experience
There isn’t a single, well-marked path from designer to GM. You can start building the experience by taking responsibility for a problem that reaches beyond design.
Operate with high agency Pick a problem affecting customers, even if it isn’t on the design roadmap. Bring the right people together, propose a next step, and agree on how you’ll measure progress. High agency includes making the follow-through your responsibility: checking the results, surfacing what isn’t working, and helping the team adjust.
Build relationships with customers and partners Spend time with customers beyond the research session. Learn why they buy, what slows adoption, and what would make them leave. Build relationships with the people who sell, support, and help deliver the product, including external partners. Those conversations connect what customers need with what it takes to serve them.
Own your area A head of design can build GM experience by setting priorities, managing a budget, developing people, and connecting the team’s work to business results. The practice starts with decisions you can own today.
Choose a scope you can be accountable for, whether that’s onboarding, a growth initiative, or a small product. Agree with your partners on the outcome and the decisions you own, then stay with it after launch.
Operate like an owner before you get the role
Nobody will give you permission to take on a GM role. You earn a shot at it through the outcomes you deliver.
Replit was where I got to go beyond design and own marketing as well, allowing me to be the DRI for the self-serve funnel. It was the amalgam of my previous career experience that enabled me to take a crack at it: performance marketing at ExactTarget, plus brand and growth at Webflow.
Have a stake in your design
Getting your designs shipped is only the beginning. Practicing extreme ownership means you care about business outcomes: customer value that leads to revenue. There is knowing the business, and then there is changing the business. Even business-oriented designers don’t all practice ownership. They may have strong business acumen, but that doesn’t guarantee they’ll push for outcomes. In many cases, they face pushback from cross-functional stakeholders.
As I tell my heads of design, it’s time to go from appeasing stakeholders to being the stakeholder.
The GM role is often held by people with MBAs, people in product, and operators. This was once true of the product manager role. We are now seeing more designers and people from other disciplines become product managers, and they can become GMs too.
The Designer's Toolkit for the AI Age (Website)
A growing design pattern at the intersection of AI and systems allows interfaces to render live in chat rather than just being described.
Summary
Decoder
- Real-Time UI: Interfaces generated and rendered live by an AI model during a conversation rather than being retrieved from a pre-built component library.
Original Article
Full article content is not available for inline reading.
Apple leak claims 20th-anniversary iPhone will ape a design formula from a decade ago
Apple is reportedly testing a display redesign for its 2027 iPhone Pro models, featuring larger, quad-curved, nearly borderless screens.
Summary
Original Article
Apple is reportedly testing a major display redesign for its rumored 20th-anniversary iPhone Pro models, featuring larger quad-curved, nearly borderless screens with a smaller display cutout. While the current iPhone 18 Pro focuses on internal upgrades, the 2027 models could deliver a more dramatic visual refresh—though the leak reflects early prototypes, and the design may still change.
The mascots, design and visual history of Wikipedia, the encyclopedia that puts functionality above aesthetics
Wikipedia remains a bastion of data-first design, resisting the attention economy by prioritizing functional simplicity and volunteer-driven aesthetics over modern UI trends.
Summary
Decoder
- Vector 2022: The latest major design update to the Wikipedia desktop interface, criticized by some for being too subtle, intended to improve readability and consistency.
- Wikimedia Commons: A massive repository of freely usable media files that contributes to the visual ecosystem of all Wikimedia projects.
Original Article
Wikipedia's enduring design is built around simplicity, accessibility, and trust, evolving only subtly while remaining highly customizable and supported by a passionate global community of volunteers. Beyond the encyclopedia itself, its visual identity extends through iconic branding like the puzzle globe, community-created mascots, creative projects, and campaigns that celebrate Wikipedia as a collaborative, human-powered ecosystem rather than just a website.
NEW STANDARD.S gives climate journalism critic Brandmelder an identity
Brandmelder adopted a minimalist, dot-grid visual identity to establish authority and neutrality as a climate journalism watchdog.
Summary
Decoder
- Climate Stripes: A visualization method showing long-term temperature changes, often used as a color coding system in environmental data design.
Original Article
Brandmelder's new identity was designed to position the climate journalism watchdog as a credible, neutral media critic rather than an activist organization. Built around a customizable dot-grid system, restrained typography, and a subtle color palette inspired by climate stripes, the branding emphasizes consistency, objectivity, and ease of use while helping the project reach mainstream news audiences.
Your Presentations Feel Lifeless Because They're Missing a Villain
Great design stories succeed not by focusing on the designer's solution, but by framing the status quo as a villain that the audience must overcome.
Summary
Original Article
The villain in a design story is never a person: it is the status quo, and it is already winning when the story opens.
Here's why Apple scrapped its shorter Apple Pencil designed for iPhone Duo: report
Apple abandoned plans for a specialized Apple Pencil for the unreleased iPhone Duo due to unresolved technical constraints.
Summary
Original Article
Apple scrapped its planned iPhone Duo-specific Apple Pencil due to technical issues.