Everyone Wants to Build AI Apps — I’m Betting on What Powers Them. Here’s Why.
Here’s why I’m betting on the picks and shovels of AI, not the apps.
Opinions expressed by Entrepreneur contributors are their own.
Key Takeaways
- Inference, not training, is the real race in AI. The race to train the biggest model has one clear winner, but the far more interesting and less crowded race is inference— running trained models fast, cheaply and at scale.
- Inference-focused chipmakers like Cerebras, Groq, Positron, d-Matrix and Tenstorrent are where the real value sits, not the app layer above them.
- The application layer will keep churning through winners and losers as models commoditize, but the infrastructure underneath it compounds quietly in the background.
Every founder I talk to has an AI application idea. Almost none of them are thinking about what those applications actually run on.
That’s the gap I’ve been investing in for the past two years — and 2026 has made the case louder than I expected.
Training was the first race. Inference is the real one.
The early AI chip narrative was all about training — who could build the biggest cluster to teach the biggest model. That race increasingly has one obvious winner, and it isn’t a startup. The far more interesting, and far less crowded, race is inference: running trained models fast, cheaply and at scale, in production, for actual paying customers. That’s an efficiency problem, not a brute-force one, and efficiency problems are where specialized architecture beats general-purpose hardware.
That’s the thesis behind my largest position, Cerebras Systems. I backed Andrew Feldman’s team in a very early round, when the company was valued at around $2 billion — well before its Series F pushed past $4 billion, before the $1 billion Series H that valued it at roughly $23 billion in February, and long before what came next.
That conviction was validated in the biggest way possible in May, when Cerebras went public on the Nasdaq under the ticker CBRS, pricing its IPO at $185 a share, raising $5.55 billion — the largest U.S. tech IPO in years — and popping 68% on debut to a market cap near $95 billion. Wafer-scale compute solved the inference latency problem in a way incremental GPU improvements couldn’t, and the public market has now agreed, emphatically.
The pivot that proved the thesis
Groq is the clearest evidence that inference, not training, is where the value is settling. Seven months after signing a $20 billion chip licensing deal with Nvidia, Groq raised another $650 million — not to keep competing on training hardware, but to double down on its inference-focused AI cloud business. A company that could have taken Nvidia’s money and exited chose instead to re-commit to inference. That’s a signal worth paying attention to.
Positron is the newest proof point. After raising $230 million in a Series B backed by Arm and the Qatar Investment Authority in February, the company’s valuation “skyrocketed” in a follow-on round announced in September, with $875 million raised to keep building inference-focused chiplet architecture aimed squarely at Nvidia’s core business. Capital is moving fast into this category because the opportunity window is real, not theoretical.
d-Matrix rounds out my thesis at the earlier stage: A $275 million Series C last November valued the company at $2 billion, built specifically around digital in-memory compute for inference workloads.
Tenstorrent, the AI chip company built by legendary architect Jim Keller, just delivered the loudest proof point yet. After raising $693 million in a Series D at roughly a $2.6 billion valuation in December 2024, the company closed two Series E tranches in August 2026 totaling about $1.4 billion, pushing its valuation to $5.76 billion. That kind of markup, this fast, only happens when the market is convinced the specialized-silicon thesis is right — and I’m glad to have been in early.
The lesson for founders, not just investors
I didn’t get into these positions because I’m a semiconductor expert. I got into them because I run an operating business that increasingly depends on AI tools, and I asked a founder’s question: Who actually captures the value when every company on earth needs cheaper, faster inference?
The answer wasn’t not only the app layer, but also the infrastructure underneath it — built by teams willing to bet on specialized architecture instead of chasing the incumbent’s playbook.
If you’re building anything AI-adjacent right now, don’t just ask which model to use. Ask who’s building the rails those models will run on for the next decade. That’s usually where the durable value — and the durable investment opportunity — actually sits.
The application layer will keep churning through winners and losers as models commoditize, but the infrastructure underneath it — the wafers, the interconnects, the inference stacks — compounds quietly in the background. I’d rather own the toll roads than bet on which car wins the race.
Disclosure: I hold investments in Cerebras, Groq, Positron, d-Matrix, and Tenstorrent, companies discussed in this article, and therefore have a financial interest in their performance. These interests represent a potential conflict of interest. The views expressed are my own and are provided for informational purposes only, not as investment advice or a recommendation to buy or sell any security.
Key Takeaways
- Inference, not training, is the real race in AI. The race to train the biggest model has one clear winner, but the far more interesting and less crowded race is inference— running trained models fast, cheaply and at scale.
- Inference-focused chipmakers like Cerebras, Groq, Positron, d-Matrix and Tenstorrent are where the real value sits, not the app layer above them.
- The application layer will keep churning through winners and losers as models commoditize, but the infrastructure underneath it compounds quietly in the background.
Every founder I talk to has an AI application idea. Almost none of them are thinking about what those applications actually run on.
That’s the gap I’ve been investing in for the past two years — and 2026 has made the case louder than I expected.
Training was the first race. Inference is the real one.
The early AI chip narrative was all about training — who could build the biggest cluster to teach the biggest model. That race increasingly has one obvious winner, and it isn’t a startup. The far more interesting, and far less crowded, race is inference: running trained models fast, cheaply and at scale, in production, for actual paying customers. That’s an efficiency problem, not a brute-force one, and efficiency problems are where specialized architecture beats general-purpose hardware.