SEMIFIVE has signed a KRW 70.3 billion ($52 million) contract with a U.S.-based AI fabless company to develop a next-generation AI inference accelerator, marking the Korean semiconductor company’s first North American “Spec Hand-off” engagement. The project comes as hyperscalers and cloud providers increasingly turn to custom silicon to optimize AI inference performance, power consumption and infrastructure economics.
The economics of AI inference are pushing cloud providers and AI companies to look beyond general-purpose accelerators and toward chips designed around specific workloads. SEMIFIVE’s latest contract illustrates how that shift is creating opportunities for custom ASIC developers that can take responsibility for chip development earlier in the design cycle.
The South Korean semiconductor company has signed a contract worth approximately KRW 70.3 billion, or $52 million, with a U.S.-based AI fabless company to develop a next-generation AI inference accelerator.
SEMIFIVE says the engagement is its first “Spec Hand-off” project in North America and its largest single contract to date. The company reported KRW 168.4 billion in total orders during 2025 and KRW 118.9 billion in new orders during the first half of 2026, meaning the new contract is equivalent to more than 40% of its 2025 orders and about 60% of its first-half 2026 new orders.
The financial scale of the project is notable, but its development model may be equally important.
Moving ASIC development upstream
Under SEMIFIVE’s Spec Hand-off model, customers provide key performance requirements and specifications while SEMIFIVE takes responsibility for detailed chip design, software development, packaging, testing and mass production.
That differs from a conventional turnkey model, where a customer typically completes much of the chip design before outsourcing manufacturing and implementation activities.
The approach effectively makes the ASIC supplier a development partner earlier in the semiconductor lifecycle. For AI companies without the internal resources to manage every stage of custom silicon development, that can reduce the number of engineering functions that need to be coordinated separately.
The model is becoming more relevant as AI workloads become more specialized.
Google and Amazon, for example, have developed custom AI silicon for their cloud infrastructure, while semiconductor companies such as Broadcom and Marvell have been involved in custom silicon programs for major technology customers. The wider industry is also seeing cloud service providers increase investment in internally designed chips as inference workloads grow.
TrendForce forecasts ASIC-based systems will account for nearly 28% of global AI server shipments in 2026, with the share potentially approaching 40% by 2030.
Designing for inference economics
SEMIFIVE’s new accelerator is being designed specifically around the requirements of large-scale AI model inference.
Inference is becoming an increasingly important workload as generative AI moves from experimentation into production. Unlike model training, which involves large periodic compute workloads, inference can run continuously as users interact with AI applications and as agents execute tasks.
Gartner forecasts that global spending on AI-optimized infrastructure as a service will reach approximately $42.3 billion in 2026, representing 96.4% growth from 2025. The firm also expects inference spending to reach $23.3 billion in 2026, surpassing the $19 billion projected for training.
That shift puts greater emphasis on performance per watt, memory bandwidth and the total cost of operating AI infrastructure.
SEMIFIVE says its accelerator will use LPDDR6, a next-generation low-power memory technology, to reduce memory-related power consumption. The design will also use PCIe Gen5 to increase data-transfer bandwidth and reduce bottlenecks between components.
For inference systems, these characteristics can be important because moving data between memory and compute resources can become a significant part of overall workload performance and energy consumption.
Big-die designs target larger AI workloads
The project will also draw on SEMIFIVE’s experience with large-area chip designs. The company says it has worked on multiple “Big Die” projects involving dies of up to 800 square millimeters.
The new accelerator is intended to balance high compute throughput with operational stability. Details about the accelerator’s final architecture, process node, manufacturing partner and expected performance have not been disclosed.
SEMIFIVE recently began mass production of HyperAccel’s Bertha AI inference accelerator using Samsung Foundry’s 4nm process, in a separate project involving a die larger than 500 square millimeters. That program marked SEMIFIVE’s first large-scale mass-production project using Samsung’s 4nm process.
The experience could be relevant as chip designers increasingly deal with the thermal, yield, packaging and interconnect challenges associated with larger AI silicon.
Custom silicon expands beyond hyperscalers
The growth of custom AI accelerators is not necessarily about replacing GPUs across every workload. Instead, specialized ASICs can be used where a particular workload is stable enough to justify dedicated hardware.
This is particularly relevant for hyperscalers and cloud service providers operating AI services at large scale. Once a model architecture or inference workload reaches sufficient volume, even relatively small improvements in power consumption, throughput or utilization can translate into substantial infrastructure savings.
Gartner forecasts worldwide AI spending will reach $2.7 trillion in 2026, up 49.5% year over year. AI infrastructure—including AI-optimized servers, networking, processing semiconductors and related systems—remains the largest area of spending growth.
The semiconductor industry is consequently becoming an increasingly important part of AI platform strategy. Gartner forecasts AI processing semiconductor revenue will grow at a 26.8% compound annual growth rate through 2030.
That growth is also changing the competitive landscape. NVIDIA remains a dominant supplier of AI accelerators, while AMD, Google, Amazon and other technology companies are developing alternative architectures. Gartner says the shift toward inference-optimized architectures and greater supply-chain diversification is challenging established approaches in the AI semiconductor market.
From chip design to AI infrastructure
SEMIFIVE’s contract illustrates another development in the AI semiconductor market: custom silicon providers are increasingly participating in the system design process rather than simply implementing customer-created chip specifications.
The company’s Spec Hand-off model places requirements definition, architecture, software and manufacturing coordination within a broader development engagement.
Following a planned tape-out in the first half of 2027, SEMIFIVE and its customer expect global mass production to begin in 2028, targeting hyperscalers and cloud service providers.
The timeline reflects the long development cycle of custom AI silicon. By the time the accelerator reaches production, inference workloads may have evolved significantly from today’s systems. That makes architecture flexibility, software compatibility and efficient memory and interconnect design as important as raw compute capacity.
For SEMIFIVE, the North American contract expands its role in a market increasingly defined by specialized AI infrastructure. For the broader semiconductor industry, it is another indication that the next phase of AI computing will involve a mix of GPUs, custom ASICs, specialized accelerators, memory technologies and high-speed interconnects rather than a single hardware architecture.
Market Landscape
The AI semiconductor market is shifting as inference becomes a larger share of production AI workloads. Gartner expects inference spending to exceed training spending in 2026, while TrendForce projects ASIC-based AI servers will represent nearly 28% of AI server shipments this year.
Hyperscalers are investing in proprietary silicon to optimize workloads, diversify accelerator supply and improve infrastructure economics. Google and Amazon have developed their own AI accelerators, while companies such as Broadcom and Marvell participate in custom silicon programs.
Gartner expects AI processing semiconductor revenue to grow at a 26.8% CAGR through 2030, underscoring the expanding market for specialized AI compute.
Top Insights
- SEMIFIVE signed a $52 million contract to develop a next-generation AI inference accelerator for a U.S. fabless AI company.
- The project is SEMIFIVE’s first North American Spec Hand-off engagement and its largest single contract to date.
- LPDDR6 and PCIe Gen5 are intended to improve memory efficiency and high-volume data movement for inference workloads.
- TrendForce forecasts ASIC-based systems will account for nearly 28% of AI server shipments in 2026.
- Gartner expects global AI inference spending to reach $23.3 billion in 2026, surpassing projected training expenditure.
Power Tomorrow’s Intelligence — Build It with TechEdgeAI











