[For more diverse corporate information on the startups mentioned in this article, please visit the Unicorn Factory Big Data Platform 'Data Lab'.]

#On the 23rd of last month (local time), a surprise guest appeared at AMD's annual conference "Advancing AI 2026" held in San Francisco, USA. The keynote speaker was Andrew Feldman, CEO of Cerebras, an inference semiconductor developer introduced directly by Lisa Su, AMD CEO.
At the event, CEO Su and CEO Feldman announced that Cerebras' "CS-3" would be integrated with AMD's "Helios." Helios is an AI server rack equipped with AMD's top-tier semiconductors. The plan is to drastically boost AI inference speeds by integrating Cerebras' CS-3, which offers AI inference performance more than 20 times faster than GPUs (graphics processing units).
#Last December, a similar move was made by NVIDIA. NVIDIA acquired the assets and key personnel of Groq, an inference semiconductor startup, for $20 billion (29 trillion won). Jensen Huang, NVIDIA CEO, emphasized at GTC 2026 in March this year that Groq's technology would be integrated into the top-tier AI server rack "Vera Rubin" to enhance AI inference speeds.

Global semiconductor giants are increasingly reaching out to inference semiconductor startups. In addition to NVIDIA and AMD, Intel announced cooperation plans with SambaNova, an inference semiconductor startup, earlier this year. A few weeks after announcing its collaboration with Cerebras, AMD also declared its intention to acquire another inference semiconductor startup, Talas.
In response to these developments, domestic semiconductor startups such as Rebellion and │Furiosa AI are paying close attention. The NPUs (neural processing units) they develop are also inference semiconductors. These companies evaluate that this series of moves will help dispel doubts about whether a market for inference semiconductors truly exists beyond GPUs.
A Rebellion representative stated, "We view this as a positive signal acknowledging the demand for inference semiconductors in the big tech industry," while a Furiosa AI representative added, "It is significant that heterogeneous AI infrastructure is emerging as the market trend."

However, not all views are positive. There are concerns that market entry may become difficult if inference semiconductors are not partnered with big tech companies like Groq and Cerebras.
In response, industry experts argue that the direction of inference semiconductors aligning with big tech differs from domestic NPUs. The key feature of Groq and Cerebras chips is their extremely fast speed. They use SRAM (static random-access memory) with small capacity but wide bandwidth. Their stance is that problems arising from small capacity can be resolved by purchasing and installing more chips at a higher cost.
As a result, it is reported that Cerebras chips offer inference computation speeds 20 times faster than NVIDIA GPUs. In response, big tech companies adopt a strategy of dividing inference computations into "prefill" and "token generation (decode)" stages when collaborating with them, passing only token generation to inference semiconductors.
In contrast, domestic AI semiconductor companies' NPUs focus more on power efficiency than speed. They use HBM (high-bandwidth memory), which is fast and has large capacity, reducing the total number of chips required from the perspective of data center operators. This is a strategy to reduce capital expenditure (CAPEX). While overseas companies create expensive "supercars" that do not consider fuel efficiency, domestic companies are building "electric cars" with low maintenance costs as their strength.
A Furiosa AI representative stated, "Since our direction differs from US inference semiconductors, we do not feel significantly threatened." In fact, recently a data center in Sweden decided to install Furiosa AI's NPU alongside GPUs. This choice reflects the European data centers' emphasis on power efficiency.

Ultimately, it is expected that the AI infrastructure market will diversify according to its use cases: "absolute computation speed" versus "power and cost efficiency." Areas requiring ultra-fast computations will be dominated by big tech alliances, while domestic NPU companies will target data center areas where power savings and operational efficiency are urgent.
A deep-tech investment reviewer stated, "It has been proven that specialized inference semiconductors are needed alongside general-purpose GPUs with high versatility in AI computation." However, he added, "The market has not yet determined whether faster computation speeds or lower operating costs are more necessary."
He further noted, "For domestic NPU companies, the key will be how much demand exists in the market for their strengths in power and cost efficiency."
[MoneyToday Startup Media Platform Unicorn Factory]