Corvex has launched Token Factory, a serverless AI inference platform designed to give developers and enterprises access to open-weight models without operating their own GPU infrastructure. The platform runs models on Corvex-managed hardware, with zero data retention enabled by default and compatibility with OpenAI- and Anthropic-style APIs.
Running open-weight AI models can give enterprises more control over model selection and data, but operating the underlying GPU infrastructure remains a significant technical and financial challenge. Corvex is targeting that gap with Token Factory, a serverless inference platform that provides API access to open-weight models without requiring customers to deploy or manage GPU clusters.
The company says Token Factory initially supports GLM 5.3 from Z.ai and DeepSeek V4 Flash 0731. Instead of sending inference requests to model developers or third-party inference providers, Corvex says it operates the models itself on hardware it manages.
That architecture is central to the company’s pitch around data control. Corvex says prompts and responses are processed in memory and, by default, are not logged, stored or used to train models. The company does retain limited operational metadata for security, service operations and billing, but says that metadata does not contain prompt or response content.
The distinction is becoming increasingly important as businesses move generative AI beyond experimentation. Enterprise AI applications may process source code, customer records, internal documents and other proprietary information, making the location and handling of inference data a consideration alongside model performance and cost.
Token Factory is designed to reduce the infrastructure burden associated with running open-weight models. Developers pay for the input and output tokens they consume rather than maintaining GPU capacity that may sit idle between workloads.
Corvex describes the service around four priorities: data assurance, reliability, performance and simplicity. Its API compatibility with OpenAI and Anthropic interfaces is intended to allow developers to connect existing AI clients and coding tools with relatively few configuration changes.
In practical terms, developers can change the API endpoint, authentication key and model selection while leaving much of an existing application workflow intact. Corvex cautions that feature compatibility can vary depending on the client and model.
The approach puts Token Factory into an increasingly competitive AI inference market. Cloud providers including Amazon Web Services, Microsoft Azure and Google Cloud offer managed access to foundation models, while specialist inference providers compete on model availability, latency, throughput and pricing.
Open-weight models introduce another dimension to that competition. Organizations can select models outside the proprietary ecosystems operated by the largest AI companies and, depending on the model’s license and deployment requirements, gain more flexibility over how those models are hosted.
However, open-weight does not automatically mean simple. Running models efficiently requires GPU capacity, memory management, model optimization, networking and ongoing infrastructure operations. Inference providers can abstract much of that complexity, turning model execution into an API rather than an infrastructure project.
Corvex is attempting to combine that abstraction with a security-oriented deployment model. The company says it does not route requests to third-party inference providers, which could reduce the number of infrastructure layers through which sensitive prompts and responses travel.
The company’s SOC 2 Type II certification provides another enterprise security credential, while Corvex says it can sign Business Associate Agreements with customers subject to HIPAA. Additional security documentation is available through its Trust Center.
Performance will be another important factor. Inference platforms need to balance time to first token with sustained throughput, particularly for interactive applications and AI-agent workloads. Corvex says it tunes its inference stack to manage both measures under load, although real-world performance will depend on the selected model, workload and deployment configuration.
The company is also positioning Token Factory as an entry point into more dedicated infrastructure. Its roadmap includes dedicated enterprise deployments and private inference environments, suggesting a progression from shared serverless inference toward more controlled architectures for organizations with stricter security, compliance or performance requirements.
That progression mirrors a broader enterprise AI infrastructure trend. Organizations are increasingly evaluating not just which model performs best, but where inference runs, who controls the infrastructure, how data is retained and how costs scale with usage.
For Corvex, the opportunity is to make open-weight models easier to consume while preserving some of the control that makes self-hosting attractive. The challenge will be delivering competitive model performance and economics without passing the operational complexity of GPU infrastructure back to customers.
Token Factory’s launch therefore represents more than another model API. It reflects the continuing shift in AI infrastructure toward managed inference, where developers can choose open-weight models while outsourcing the hardware and systems engineering required to run them in production.
Market Landscape
The AI inference market is moving toward managed, serverless and specialized model execution, giving developers access to increasingly capable models without requiring direct ownership of GPU infrastructure.
Corvex is positioning Token Factory between traditional cloud infrastructure and fully proprietary AI platforms. Its focus is open-weight models, API compatibility and data-control requirements, while its infrastructure handles model execution.
Competitors include hyperscale cloud platforms, specialist inference providers and GPU infrastructure companies. Differentiation increasingly depends on latency, throughput, cost per token, model availability, privacy controls and deployment flexibility.
For enterprises, the choice is becoming less about simply selecting a model and more about selecting an inference architecture that balances performance, security, compliance and operational complexity.
Top Insights
- Corvex Token Factory provides API access to open-weight AI models without requiring customers to operate their own GPU clusters.
- Prompts and responses are processed in memory, with zero data retention enabled by default and no use of customer content for model training.
- OpenAI- and Anthropic-compatible APIs are designed to simplify migration from existing AI applications and development tools.
- The platform charges for consumed input and output tokens rather than idle GPU capacity, creating a usage-based inference model.
- Corvex plans dedicated deployments and private inference environments for enterprises requiring greater infrastructure control.
Power Tomorrow’s Intelligence — Build It with TechEdgeAI
