NewsStocksAolani Launches Token Factory for Managed AI Inference in Asia

Aolani Launches Token Factory for Managed AI Inference in Asia

Author: Citybuzz·

Key Takeaways

  • Aolani Token Factory uses a per-token billing model, with customers prebuying credits and paying based on token usage.
  • Aolani manages the full inference stack, including GPU allocation, model serving, orchestration, scheduling, and workload optimization.
  • The platform launches with support for open-source models including DeepSeek, GLM, Kimi, and Qwen, and it can also host custom models through OpenAI-compatible APIs.
  • Dedicated capacity and data isolation options are available for enterprises with compliance and data residency requirements.
  • Aolani says the platform is aimed at AI agents, enterprise copilots and knowledge assistants, and coding agents for development tasks.
Aolani Launches Token Factory for Managed AI Inference in Asia

Singapore-based neocloud Aolani has unveiled the Aolani Token Factory, a managed inference platform designed to let organizations deploy and scale AI models on a pay-per-token basis, eliminating the need to provision or manage underlying GPU infrastructure. The announcement positions Aolani as the first Singapore-founded neocloud to offer production-grade, managed inference at scale.

As global AI companies expand operations in Singapore and enterprises worldwide invest in AI for tangible business outcomes, demand for production-grade inference infrastructure is accelerating. That shift matters because model training has drawn much of the early attention in AI, but running models reliably in production is where organizations face recurring compute, compliance, and operational demands. The Aolani Token Factory aims to close the accessibility gap, providing AI-native companies and enterprises a compliant, high-performance path from experimentation to production deployment.

The platform operates on a per-token metering model, where customers prepurchase credits and pay based on token consumption, avoiding capital-intensive GPU investments. Aolani manages the full inference stack, including GPU capacity allocation, model serving, orchestration, scheduling, and workload optimization, allowing customers to scale consumption without continuous infrastructure provisioning.

At launch, the Token Factory supports leading open-source models such as DeepSeek, GLM, Kimi, and Qwen, with plans to expand the catalog based on customer demand. Customers can also deploy custom models via OpenAI-compatible APIs. For enterprises with strict compliance and data residency needs, dedicated capacity and data isolation options are available.

The platform targets three core production use cases: AI agents for high-volume inference and workflow automation, enterprise AI applications such as internal copilots and knowledge assistants, and coding agents for code generation, completion, testing, and review.

Sea Xu, Applied AI Research Lead at Aolani, highlighted the platform’s performance and flexibility: “The Aolani Token Factory is built on a high-performance inference stack that supports the most in-demand open-source model families. We designed the platform for fast model adaptation and deployment, so our customers can get access quickly as new models emerge. As Southeast Asia’s AI ecosystem evolves and grows rapidly, it is our goal to ensure that the infrastructure serving it keeps pace.”

Nicholas Chia, CEO of Aolani, emphasized the transformative impact: “Fast-moving AI natives want to build and ship products flexibly and on-demand, without the need to manage GPU fleets. With our competitive per-token pricing and a fully managed stack, companies can go from model selection to production deployment without the capital outlay or operational complexity of self-managed infrastructure. This is a significant milestone for us and our customers as the Aolani Token Factory will fundamentally change how customers access AI compute.”

The launch comes amid growing regional demand for AI infrastructure, including in Singapore, where companies are building and deploying more AI services that need dependable inference rather than one-off experimentation. By offering a managed, pay-as-you-go model, Aolani aims to lower barriers for businesses of all sizes, enabling them to leverage advanced AI capabilities without upfront investments. That setup may be especially relevant for startups and mid-sized enterprises that need access to production infrastructure but prefer not to build and maintain their own GPU stack.

Interested parties can register interest at Aolani Token Factory. For more information about Aolani, visit Aolani’s website.