After huge revenue beat for Q2 earnings, Nvidia buys Hugging Face

Hugging Face had been looking to raise money and was considering a sale lately. (Picture: generated)
Nvidia’s revenues were up 106% to $96.22 billion last quarter, while predicting a 70% increase in the next fiscal year, so agreeing to acquire Hugging Face for $12.9 billion might seem like pocket change.

It is, however, one of Nvidia’s largest acquisitions, Reuters notes, after the buyout was first reported by The Information.

Hugging Face is the «face» of the open source AI movement and maintains a repository of almost all available models, making this a significant infrastructure investment. AI labs treat publishing their weights on it as their official release.

Nvidia is of course not new to Hugging Face or open source, having as good as bought the OSS AI lab Poolside earlier this week to build their own models. They also invested in a $235 million funding round for Hugging Face in 2023 that valued it at $4.5 billion, and tried to invest $500 million in 2025 at a $7 billion valuation, according to The Financial Times.

It appears that Nvidia is somewhat hedging their bets and is increasingly investing in open source, as the frontier AI labs are increasingly developing their own chips, Reuters writes. Nvidia remains a significant investor in closed source providers, though.

Read more: Reuters, The Information (paywalled), and Business Insider. Discussion on r/Singularity and Hacker News.

OpenAI’s Jalapeño beats Nvidia’s Blackwell in performance per watt

The new chip will be deployed at scale within the year. (Picture: OpenAI)
Using the measure of performance per watt rather than throughput per second, OpenAI ran three open source models, GPT-OSS, DeepSeek R1 and Kimi K2.5 1T, through their new processor on the InferenceX benchmark from SemiAnalysis.

They found that the processor is wicked fast compared with «leading chips,» which means Nvidia’s offerings, specifically the Grace Blackwell 200 which they show getting trounced in some tests.

The benchmark found that Jalapeño delivers 1.9x more tokens per second when measured per watt on peak efficiency, and is 17.8x faster on higher token density, which translates to raw performance in handling requests.

They also found that end-to-end latency was between 1.7x and 3.4x lower depending on the model, meaning the user will spend less time waiting for responses.

These two measures combined are key, as most chips have to make tradeoffs between latency and throughput, and few can be good at both, OpenAI hardware vice president Richard Ho tells The Verge.

Jalapeño was developed in record time — 9 to 16 months — assisted by AI, and OpenAI says their upcoming Astra model is already busy working on the second generation, which they say is «in deep development.»

OpenAI will deploy and «operate Jalapeño at scale» within their compute infrastructure «by the end of the year.»

Read more: OpenAI’s report, X post. The Verge, Bloomberg (paywalled) and The Register. Discussion on Hacker News and r/Singularity.

Nvidia’s inference rack Groq 3 LPX delivers record 3,431 tokens per second

Nothing chews through inference tasks as fast as Nvidia’s Groq. By far. (Picture: Nvidia/generated)
The never-before-seen feat was achieved running Google’s Gemma 4 31B through Artificial Analysis’ standard tests, and is way ahead of anything on the market.

The only comparable score is that of OpenAI’s GPT-Sol running on Cerebras chips, which achieved 750 tokens per second earlier in August. They called this «Ultrafast mode.»

For more normal hardware setups, Opus 5 gets 58.8 tokens per second, and GPT-5.6-Sol clocks in at 74.4 per second, while the «faster» Gemini 3.7 Flash gets 371.1 throughput tokens.

Nvidia further says that it achieved this output score while maintaining hundreds of thousands of context tokens, and that Groq is some 34X faster than today’s quickest hardware in time to generate 5,000 tokens.

A Groq chip pairs 500 MB of high speed SRAM on die, directly next to the chip, that delivers 150 TB/s throughput. They come in racks of 256 chips stacked together for a total of 40 petabytes of memory bandwidth.

While they are great for inference tasks, GPUs will still be the workhorse of AI data centers, as they can handle both training and later inference — but Jensen Huang of Nvidia recommends setting aside 25% of data center space for the new Groq chip racks, according to CNBC.

The new test scores come as Nvidia is announcing that Groq 3 chips are now in production with Samsung and are generally available.

Read more: Nvidia’s announcement, production note. Writeups on CNBC and The Register.

Stripe buys OpenRouter, reportedly for $8 billion

OpenRouter says they are vendor neutral and builds on the same principles as Stripe. (Picture: OpenRouter, generated)
OpenRouter will continue as normal after the deal, supporting token- and task-optimized routing between models and keeping their neutrality.

Stripe is a huge fintech darling that routes some $1.9 trillion in payments each year and has a revenue of $5.1 billion from clients like Ford, Spotify, OpenAI and Anthropic. The still privately held company has attracted early investments from Elon Musk and Peter Thiel, according to Wikipedia, and they are also trying to buy PayPal, Axios reports.

OpenRouter had raised $164 million in May 2026 at a valuation of $1.3 billion, so the Axios report of an $8 billion price is at a premium as well as highlighting their stellar growth since their establishment in 2023, just as the AI market started emerging.

— Tokens are the central currency for companies building with AI, and it’s clear that the real-world economic potential will depend on making good use of scarce compute resources, says Patrick Collison, co-founder and CEO of Stripe in their press release, hinting at how price-sensitive enterprise AI usage has become.

Earlier this year, Stripe were giddish on AI, declaring that the start of 2026 also marked «the beginning of the singularity,» according to an investor letter published by Eric Newcomer, as they reached customers in 88% of the Forbes AI 50 list, according to Axios.

Read more: Stripe presser, OpenRouter release. Writeups on Axios, CNBC, and TechCrunch. Discussion on Hacker News and r/Singularity

OpenAI to lease massive 8GW compute from PORTS-Pike facility in Ohio

The new data center will be at full capacity in six years, greatly expanding OpenAI’s capacity. (Picture: generated)
The operation will be owned by SoftBank’s SB Energy and backstopped by Nvidia guarantees of $105 billion, with the first 800 megawatts of capacity coming online in 2028.

The completed capacity by 2032 would quadruple the compute currently held by OpenAI, said to be about 1.9 GW in 2025, yet growing exponentially.

The facility, built on land previously used for uranium enrichment by the US government, will exclusively use chips and networking from Nvidia, who expects it to represent 1.5 million of their GPUs and revenue of $150—$200 billion, not including future upgrades, according to Jensen Huang.

OpenAI says the data center will use its own energy and recycle its water supply, lessening the load on consumer facilities.

It will also supply the local community with 35,000 construction jobs over six years of building, and 2,500 jobs in operations once it’s finished.

In addition, OpenAI are providing $85 million in Codex credits to Ohio college students for the 2026/27 academic year, in addition to a «community grant fund» of some $40 million.

Read more: OpenAI’s announcement and X post, Nvidia’s presser, Jensen’s X post. Writeups on Axios, CNBC, and Reuters.

Meta’s AI fallback plan: Forming a cloud computing company

Meta could soon be renting out excess compute, as it continues to build out capacity. (Picture: Adobe)
Already sitting on one of the world’s largest compute clusters with plans to add some $145 billion worth more just this year — Meta is making contingency plans.

Bloomberg (paywalled) is reporting that they are now planning to form a cloud computing company to compete with AWS, Azure and Google Cloud, aiming to sell «excess capacity» to paying customers.

The idea is a hedge against overcapacity, Axios reports, quoting Zuckerberg as saying in May: «that is an option that we have, and that is partially what gives us confidence in investing in building this out.»

The new unit would offer raw compute to companies looking to train models, or could offer access to models from other AI labs and rent out the infrastructure to power them. This would be similar to services from Amazon and Azure.

SpaceX, now incorporating X.ai, faced a similar problem when their compute power got bigger than their needs, and struck deals with Anthropic and Google to sell excess compute access, Reuters notes.

It is estimated that Meta currently sits on 20 gigawatts of capacity, and it plans to add another 14 GW over the next years, Axios says.

Read more: Bloomberg (paywalled), Axios, Reuters, CNBC, and Engadget.

OpenAI delivers Jalapeño, a state-of-the-art chip made in only nine months

The Jalapeño chip offers substantially better performance per watt than anything on the market. (Picture: OpenAI)
The new inference chip, developed with help from OpenAI’s models in collaboration with Broadcom, has compressed what is normally a glacial, years-long process into just a short sprint.

— We believe [this] to be the fastest ASIC development cycle ever achieved in high-performance advanced semiconductors, OpenAI says.

The chip is designed specifically for OpenAI’s needs on current and upcoming language models, and to combine the power of current AI accelerators with the shorter latency of «specialized systems,» like Cerebras’ hardware.

— The world is moving to a compute-powered economy, says Greg Brockman, President and Co-Founder of OpenAI. — Jalapeño is part of our long-term full-stack infrastructure strategy to make compute more abundant.

The new silicon will be deployed at gigawatt-scale data centers just as fast as it was made, starting with Microsoft «and other partners» already in 2026, Broadcom says. It is only the first design in what the companies expect to be generations of products.

Jalapeño is already running LLM workloads in the lab, delivering «substantially better» performance per watt than «the current state-of-the-art,» OpenAI says.

Read more: OpenAI’s presentation, CNBC, CNN, Reuters, and TechCrunch.

The EU plans for «technological sovereignty» with huge investments

The European Union has a long checklist of things to improve in the AI age, and stands ready to invest «at scale.» (Picture: Shutterstock)
The EU is increasingly concerned at their reliance on the USA for all things cloud, software and AI, and is taking urgent steps to counter it, or, as they put it, to «strengthen Europe’s digital resilience.»

— We cannot afford to depend on others for the technologies that keep our hospitals running, our energy grids stable and our services secure, Commission President, Ursula von der Leyen says in a statement.

Continue reading “The EU plans for «technological sovereignty» with huge investments”

SoftBank to spend €75 billion on data centers, manufacturing in France

The manufacturing and data centers will require next-gen worker skills. (Picture: Adobe)
The northern port of Dunkirk will become an advanced, automated AI manufacturing hub in the new plan — to support the buildout of 5 GW of compute power.

For the first phase, SoftBank is offering €45 billion to build 3.1 GW of capacity at various sites in the Hauts-de-France region by 2031, ramping up at a later date.

— SoftBank Group’s AI data centers will support growing demand for high-performance computing from AI companies, cloud providers, enterprises, public institutions and research organizations, they write in their release.

France was selected because of its wide manufacturing base, but another important factor is it advanced power grid, fueled by generous amounts of nuclear energy.

SoftBank will partner with Schneider Electric on the project, which will also create «thousands» of advanced, high-skilled jobs, requiring local universities, engineering schools, and training institutions to step up.

SoftBank is a major investor in OpenAI and is deeply involved in the Stargate program to build out compute capacity. This effort seems separate from both.

Read more: SoftBanks release, Reuters, and TechCrunch.

Anthropic grew 80-fold last quarter, secures compute deal with SpaceX

Power hungry: Anthropic’s growth has been off the charts last quarter. (Picture: generated)
The AI lab was preparing for its usual 10x growth rate this quarter, but the numbers made a huge, unexpected jump, CNBC reports

— That is the reason we have had difficulties with compute, says CEO Dario Amodei, revealing their strain on infrastructure since at least April.

The company has now made a deal with SpaceX for instant access to their entire Colossus 1 supercomputer, consisting of some 220,000 Nvidia GPUs, totaling more than 300 megawatts.

Using all of that capacity should immediately improve access for Anthropic’s users, and they are already doubling rate limits for the Pro and Max plans, and is «considerably» increasing API rates, while removing «peak hours» restrictions.

Elon Musk says he approved the deal after evaluating Anthropic’s altruism, that xAI will be known as SpaceXAI and that they have already moved Grok training to the Colossus 2 cluster.

Read more: SpaceX announcement, Anthropic’s announcement, Anthropic on X, CNBC, Engadget.

Citing strain on servers, Anthropic secures 5 GW of compute from Amazon

It’s a deal where it appears both sides win. (Picture: Amazon)
Anthropic admits to taxing its servers lately, saying that their recent growth «places an inevitable strain on our infrastructure; our unprecedented consumer growth, in particular, has impacted reliability and performance […] especially during peak hours»

The AI lab says 1 gigawatt of the new capacity on Amazon’s custom silicon will come online in late 2026, and is committing to spending $100 billion on Amazon over the next decade to «train and run Claude.»

The deal also includes an initial investment of some $5 billion from Amazon which will scale to $25 billion «in the future,» tied to «commercial milestones.»

Andy Jassy, CEO of Amazon, has taken some flak for their gigantic $200 billion spend on AI infrastructure, but with this deal, they are recouping half of that:

—Anthropic’s commitment to run its large language models on AWS Trainium for the next decade reflects the progress we’ve made together on custom silicon, Jassy says.

Read more: Anthropic’s announcement, Amazon’s release. Writeups on CNBC, TechCrunch, and Engadget.

Amazon CEO explains $200 billion AI spend, has over $15 billion in AI revenue

AI bets are already starting to pay off for Amazon, Jassy says. (Picture: Shutterstock)
Andy Jassy’s annual shareholder letter this year was all about defending their massive, 60% increase in capital spending on AI infrastructure.

— We’re not investing approximately $200 billion in capex in 2026 on a hunch, he writes according to CNBC. — We’re not going to be conservative in how we play this — we’re investing to be the meaningful leader, and our future business, operating income, and [cash flow] will be much larger because of it, he continues.

There are already deals incoming to support this claim, he writes, noting a $100 billion commitment from OpenAI for AI compute on their server farms.

AI server use at Amazon’s cloud unit has reached more than $15 billion in annualized revenue, Jassy writes, and now represents about 10% of the total for AWS, Reuters reports.

Read more: The shareholder letter (long), writeups on CNBC and Reuters.

Anthropic reaches $30B revenue, gets compute from Google and Broadcom

Anthropic continues to diversify its compute needs. (Picutre: Anthropic)
Anthropic now says it has a run-rate revenue of $30 billion, up from $14 billion in February during their last fundraising.

They are also announcing that they are brining in new compute capacity, based on next generation Google TPUs that will start coming online in 2027.

The companies offer no detail the cost of the «partnership» or how much compute they are actually buying, but Broadcom Is hinting it’s around 3.5 GW, according to CNBC.

Anthropic also say they have doubled the rate of customers spending more than $1 million per year to 1,000, in just two months.

Claude now runs on Amazon’s Trainium chips, Google TPUs and Nvidia GPUs. The latter are more used, and Amazon remains their primary cloud provider, Anthropic says.

Read more: Anthropic’s announcement, CNBC adds numbers.

Mistral raises $830 million in debt to build data center just outside Paris

It’s a big investment for European AI, but significantly lower than what US AI labs are spending. (Picture: Mistral)
The new 44 MW data center will be powered by 13,800 Nvidia chips, and should be online by the second quarter of 2026.

The French AI lab is hoping to secure 200 megawatts of capacity by the end of 2027, Reuters reports.

— Scaling our infrastructure ​in Europe is ⁠critical to empower our customers and to ensure AI innovation and autonomy remain at the heart ​of Europe, says Mistral CEO Arthur Mensch.

The news comes hot on the heels of Mistral’s February startup of a €1.2 billion data center in Sweden, according to CNBC.

Mistral is the largest European AI lab, has contracts with the French armed forces, and has secured $3.1 billion in funding so far, TechCrunch writes.

Read more: Reuters, CNBC and TechCrunch.

Google’s next data center in Minnesota will have the world’s largest battery

Google’s energy in Minnesota won’t lead to higher electricity costs for consumers, they say. (Picture: Google)
The data center in Pine Island will have 1.9 gigawatts of capacity, sourced from new wind and solar power from Xcel Energy.

This will then be attached to a 300 megawatt iron-air battery from Form Energy, ensuring continuous service to operations.

That will be the world’s largest commercially deployed battery, and can provide power to the data center for a whopping 100 hours, TechCrunch notes.

As part of their buildout, Google is announcing that they will pay for their electricity in full, and will also invest $50 million in Xcel Energy’s green energy program to place batteries across their grid.

— Google’s partnership with Xcel Energy reimagines how data centers can be served, Google writes.

Read more: Google’s announcement, Xcel’s presser, CNBC and TechCrunch.