M

The DeepSeek Counterstrike: When American Companies Started Voting for Chinese AI with Their Wallets

In the summer of 2026, American enterprises are answering 'Is Chinese AI any good?' in the most primal way — with their wallets. From a San Francisco startup routing 100% traffic to DeepSeek, to it topping Ramp's fastest-growing paid services for the first time, this global counterstrike reveals a trillion-dollar cost awakening, a technological breakthrough behind the chip blockade, and the birth of two parallel AI ecosystems.

#DeepSeek#China AI#artificial intelligence#Huawei Ascend#chip blockade#AI cost#tech出海#modern China
15 min read
Intermediate
2026/06/28
The DeepSeek Counterstrike: When American Companies Started Voting for Chinese AI with Their Wallets

The DeepSeek Counterstrike: When American Companies Started Voting for Chinese AI with Their Wallets

AI technology abstract visual merging DeepSeek logo and chip patterns

Late one night in early June 2026, Flo Crivello, CEO of San Francisco AI startup Lindy, stared at his financial dashboard and felt a chill run down his spine. The company’s monthly AI model API bill had just surpassed the combined salaries of all 25 employees.

“This isn’t a cost problem,” Crivello told CNBC. “This is a survival problem.”

He made a radical decision: route 100% of Lindy’s AI traffic to a Chinese company — DeepSeek. After the switch, Crivello posted the company’s AI cost curve on social media. The shape of that line needed no caption: it plunged from a mountain peak straight off a cliff.

Lindy is not alone. In the summer of 2026, American enterprises are answering the question “Is Chinese AI any good?” in the most primal way possible — with their wallets.

The trillion-dollar AI bill shock

Rewind one month. Ramp, a financial platform managing expenses for over 50,000 American companies, published its June software trends report. One name stood out in a way nobody expected: DeepSeek had, for the first time, topped the “fastest-growing paid service” list. It was the first time a Chinese AI model had ever appeared on that ranking — and the first time American companies were paying a Chinese firm directly, at scale, through its official API.

This is not tech enthusiasts tasting something new. This is CFOs fighting for their lives in quarterly budget meetings.

The American tech industry is going through an unprecedented “AI bill shock.” According to data compiled by Wall Street CN from multiple sources, cumulative US corporate AI spending has now surpassed one trillion dollars — with ROI nowhere near keeping pace. Uber burned through its entire 2026 token budget in just four months. Salesforce is paying Anthropic roughly $300 million this year alone. Amazon quietly scrapped its internal AI leaderboards — employees had been inflating usage to game the rankings, rendering the data meaningless. Microsoft is phasing out Claude Code subscriptions for large numbers of employees, pivoting to cheaper alternatives.

“Model routing” has become the industry’s new watchword. For the past two years, the reigning strategy was “use the best model for everything” — GPT-5.5 for emails, Claude Opus for weekly reports, Gemini for spreadsheets. That kind of “cracking a nut with a sledgehammer” extravagance is now over. More and more companies are tiering their tasks: cheap models for simple work, reserving the eye-wateringly expensive ones only for the hardest reasoning jobs.

And in this great cost awakening, the biggest winner comes from China.

Ten thousand tokens for less than a cup of coffee

DeepSeek’s pricing is not a “slightly cheaper” story. It is a number that would make any CFO redo an entire budget spreadsheet.

Take DeepSeek V4-Pro: the price per million input tokens is ¥0.25 — roughly 3.4 US cents. For comparison, GPT-5.5 Pro costs $30. A seven-hundred-fold gap.

Even setting aside the most expensive comparison, DeepSeek’s per-task cost is about one-eleventh that of Claude Opus 4.7, roughly one-tenth that of GPT-5.0. On May 22 this year, DeepSeek permanently slashed V4-Pro API prices to one-quarter of their previous level — while competitors were still raising theirs.

Price opens the door. Performance keeps the customer.

On April 24, DeepSeek quietly released its V4 model — no press conference, no livestream, just a technical report pushed to GitHub. Yet the weight of that report had Silicon Valley AI engineers working through two weekends. V4 is a 1.6-trillion-parameter Mixture-of-Experts model supporting a one-million-token context window. It topped the open-source leaderboard for agentic coding tasks and closed in on frontier closed-source models in math, hard science, and competitive programming.

On June 27 — yesterday, at the time of writing — DeepSeek, jointly with Peking University, open-sourced the DSpark inference acceleration framework. Its core idea is making large models “run smarter” under high concurrency: through a semi-autoregressive generation architecture and a confidence-scheduled verification mechanism, single-user generation speed improved by 60 to 85 percent. Among the paper’s listed authors, one name stands out: Liang Wenfeng, DeepSeek’s founder — a man so low-profile he barely appears in public, now telling the world through a stream of signed papers: we are serious about building technology.

Developer debugging code in a high-performance data center

OpenRouter data corroborates this quiet revolution from another angle. OpenRouter is the world’s largest AI API aggregation platform, connecting hundreds of models to millions of developers. In the third week of May 2026, DeepSeek V4-Flash logged 3.43 trillion tokens on the platform — ranking number one globally. Add up all Chinese models, and they now account for roughly 60% of OpenRouter’s total API call volume.

Sixty percent. That is not a number anyone can ignore.

The counterattack behind the blockade

America’s chip blockade, ironically, played no small role in getting DeepSeek to this point — though the effect may be the opposite of what Washington intended.

Since October 2022, the US government has steadily tightened export controls on AI chips to China. First came the ban on A100 and H100. Then the “China-specific” A800 and H800. Then the neutered-to-the-bone H20. By the time the Biden administration’s final round of restrictions landed in late 2024, the NVIDIA chips Chinese AI companies could legally buy offered only a fraction of the performance available to their American counterparts.

Washington’s logic was simple: strangle the chips, strangle Chinese AI.

But business-world stories rarely follow the scripts of policymakers. When “can’t buy” shifted from a temporary headache to a permanent constraint, China’s tech companies made a rational collective choice: stop buying. Build it yourself.

The rise of Huawei’s Ascend chips is the critical variable in this story. According to a Goldman Sachs report cited by TMTPost, the Huawei Ascend 950PR is expected to reach mass production in the second half of 2026. On the procurement side, Ascend chips cost roughly one-quarter the price of restricted NVIDIA models. On the performance side, single-card computing power reaches approximately 2.87 times that of the restricted NVIDIA equivalents. In plain terms: more compute for less money.

Ascend’s software ecosystem is catching up fast. The CANN framework now has over four million developers, more than 3,000 partners, over 200 adapted open-source models, and 43 mainstream large models that have completed pre-training on Ascend hardware. The latest CANN Next achieves over 95% CUDA compatibility — slashing the time needed to migrate a model from NVIDIA to Ascend from weeks down to hours.

Closeup of Huawei Ascend AI processor chip

The Chinese AI industry has dubbed 2026 “Year One of domestic-chip-trained large models.” Early in the year, Zhipu AI’s GLM-Image became the first model trained entirely on domestic chips to achieve SOTA-level performance. ByteDance, Tencent, and Alibaba are now “scrambling to buy” Huawei chips — not out of patriotic sentiment, but because real-world benchmarks on DeepSeek V4 show that end-to-end inference latency on Ascend clusters is actually 35% lower than on equivalently-sized NVIDIA clusters.

NVIDIA CEO Jensen Huang said something in an internal meeting that was later widely quoted: “If DeepSeek’s latest model fully adapts on Huawei chips first, that would be a catastrophic blow to America’s strategic position in global AI.”

There is no exaggeration in that sentence — or at least, it is not hyperbolic. When the world’s most advanced AI models can run on chips not made in America, the equation of “chips equal power” begins to crack.

The birth of two AI worlds

Ramp’s chief economist Ara Kharazian added a telling caveat after the report’s release: “I wouldn’t overestimate the persistence of this trend. Direct access to DeepSeek involves real competition and security concerns.”

He is right. Data privacy, geopolitics, compliance risk — these are genuine headwinds. Not every American company is willing to send its data to servers in China. But Kharazian also acknowledged a more fundamental fact: cost pressure is not going away. And as long as the price gap exists, the market will find its direction.

What we are witnessing is not a simple “technology substitution.” It is a structural bifurcation of global AI infrastructure.

On one side, the NVIDIA ecosystem — 4.5 million CUDA developers and over 90% of the world’s AI toolchain, an unshakeable presence in the near term. On the other, an emerging Huawei Ascend ecosystem — an increasingly complete domestic AI technology stack spanning chips, frameworks, large models, and applications.

These two ecosystems are not a Cold War-style “decoupling” — too many cross-dependencies remain. But neither are they in a “one leads, one follows” relationship anymore. The more accurate description: two tracks running in parallel.

Global AI connection concept with interconnected networks

The way Chinese AI reaches the world is changing, too. It used to be “hardware exports” — chips, servers, data center equipment. Now it is “token exports” — through submarine fiber-optic cables, Chinese-trained large models flow to global developers as API services. This is not a path China actively chose; it is the path the chip blockade forced open. But it is working.

TMTPost, in a deep-dive analysis, quoted an anonymous AI entrepreneur’s lament: “We originally set out to build a global company — never expected to get blocked into becoming a ‘domestic champion.’ But looking back, the blockade cut away 90% of our speculative options. The remaining 10% are the hard bones we had to crack anyway.”

From blue jeans to tokens

Roughly twenty years ago, “Made in China” meant cheap clothing, toys, and electronic accessories in the minds of global consumers. A decade ago, it became Huawei phones, DJI drones, and TikTok. By 2026, “Made in China” has acquired an entirely new category: tokens.

DeepSeek itself does not do grand narratives. Their official technical blog is nothing but formulas and experimental data; their GitHub READMEs are as dry as academic papers. In a low-key statement after V4’s launch, they even admitted outright that “the new model’s capabilities still lag behind leading competitors by roughly three to six months” — a candor that, in an AI industry addicted to myth-making, reads as a strange sort of confidence.

But that is precisely what makes Chinese AI’s current moment so compelling. It did not win users by telling a “China story.” It won them with something that jolts a San Francisco founder awake at midnight and leads to a “100% switch” decision: cheaper, faster, good enough.

After completing the switch, Lindy’s Crivello posted a tweet so simple it was just one line: “Some decisions are hard. This one wasn’t.”

When commercial rationality begins to outweigh geopolitical narratives — when a CFO’s budget spreadsheet starts to speak louder than the White House’s export control list — the real inflection point for an industry may have already arrived, quietly.