NewsStocksHong Kong's Votee AI Builds a Cantonese Model to Challenge English and Mandarin's Dominance

Hong Kong's Votee AI Builds a Cantonese Model to Challenge English and Mandarin's Dominance

Author: Fortune Crypto·

Key Takeaways

  • Leading AI models remain strongest in English and Mandarin, which leaves many other languages with weaker support.
  • Votee AI retrains open-weight models on Cantonese data and sells them to banks, universities and government departments.
  • Cantonese differs significantly from Mandarin in grammar and vocabulary, and mainstream models often struggle with local and cultural knowledge.
  • Votee says its Cantonese data corpus has grown from 100 million tokens to more than 500 million tokens.
  • The company says it is profitable through client contracts and is in active discussions with AI Singapore while planning broader expansion in Southeast Asia.
Hong Kong's Votee AI Builds a Cantonese Model to Challenge English and Mandarin's Dominance

The AI boom is being driven by two languages: English and Mandarin Chinese. The leading AI developers—OpenAI, Anthropic, DeepSeek, Moonshot AI, Z.ai, and others—are almost all based in either the U.S. or mainland China, and as a result their models perform best in their native tongues. That leaves countless other languages and dialects, even those spoken by tens of millions of people, by the wayside.

“The whole AI revolution is in English and Mandarin,” says Pak-Sun Ting, CEO of Votee AI, a Hong Kong-based startup striving to build AI models for Cantonese and other neglected tongues. “There’s only a very small fraction that represents other languages.”

Votee is one of a growing number of companies trying to tackle AI’s neglect of other languages. It takes open-weight models from developers like Meta and Alibaba, retrains them on Cantonese data, and sells the result to banks, universities, and government departments. In practice, that matters because the users most likely to need Cantonese support are often working in public-facing settings, where accurate handling of local language can affect day-to-day operations rather than just casual conversation.

“It’s more than just culture. Cantonese is used in education, healthcare, and police communications,” Ting says. “If those don’t get covered, then AI is essentially useless.”

How Cantonese Is Different

Cantonese is often referred to as a “dialect” of Chinese, but that moniker undersells just how different it is from Mandarin, the most commonly spoken version of the Chinese language. Cantonese uses different grammar than Mandarin, as well as different vocabulary—particularly in the city of Hong Kong, where speakers often switch between English and Cantonese words, sometimes in the same sentence.

Over 80 million people speak Cantonese, roughly equal to the number who speak Korean, and more than those who speak Italian or Thai. Yet that large speaker base has not produced a comparably deep pool of standardized written data, particularly for colloquial Cantonese, which helps explain why general-purpose models can appear competent on basic prompts while still missing the nuance needed for local use.

Leading models are not completely useless at handling Cantonese. HKCanto-Eval, a set of benchmarks developed by researchers at Kyushu University, the Education University of Hong Kong, and the local AI community hon9kon9ize, and sponsored by Votee, reports that while mainstream models can handle everyday Cantonese at a reasonable level, they routinely fail when it comes to cultural and local knowledge.

Training on the Cheap

Ting says that building a Cantonese LLM is “essentially taking the same steps as if you were training a model from scratch”—taking an existing open-source model, such as Meta’s Llama or Alibaba’s Qwen, and doing additional training with Cantonese data.

Votee sources its Cantonese data from online scraping, including content from Radio Television Hong Kong (RTHK), the city’s public broadcasting service. The startup also receives data from the community and universities, taps content from its previous business as a big data company, and uses synthetic data—creating its own Cantonese data sets to train its model.

Together, these efforts have grown the corpus of Cantonese data from 100 million tokens to more than 500 million.

Votee AI’s models are around 70 billion parameters in size, significantly smaller than the best models on the market. Still, Ting says they are capable enough to understand and reason in Cantonese. More importantly, Votee’s training costs, while not trivial, remain significantly smaller than those of frontier labs. Ting estimates that the company uses between 500 million and 1 billion tokens to train its models, compared with the trillions used for English-language models, at a cost of roughly $250,000.

Sovereign AI

Several companies are working to build models for what are deemed “low resource languages,” or those that do not have a massive corpus of published work behind them. Indosat, Indonesia’s second-largest telecoms company, is building Sahabat AI, an open-source large language model that focuses on Indonesian languages like Bahasa. Singapore’s state-backed AI Singapore runs SEA-LION, covering 11 under-resourced Southeast Asian languages. South Korea has gone further, staging a state-sponsored elimination tournament—dubbed the “AI Squid Game” by local media—to pick national champions for homegrown foundation models, backed by a 2026 AI budget of roughly $6.8 billion.

All of it sits under the banner of “sovereign AI,” or the idea that governments and companies will want to own their own data, models, and infrastructure rather than renting them from overseas.

“AI has become such an essential need, and so you don’t want to be tethered to anybody else who can turn it off,” Ting says. He is candid that the full version of the sovereign AI idea—where countries own every part of the AI supply chain—is “very difficult.” Instead, he suggests countries focus on owning the foundation models and the applications built on top of them.

Governments generally do not need a model as powerful as what exists at the frontier to automate a few tasks; Ting explains that a tiny model, even one with as few as 1 billion parameters, can suit those purposes. If governments need more advanced capability, they can route a powerful English- or Chinese-language model’s outputs through a smaller, local-language output.

Ting points out that the company works with MiniMax and SenseTime models and can use Nvidia chips in its operations. “We can use Nvidia chips, we can use Moonshot or DeepSeek’s model,” he says. “We’re that person in high school who’s friends with everyone.”

For all Ting’s talk of preservation, Votee is a for-profit company. Governments and corporates are the first customers for an AI model in a language like Cantonese or Bahasa. Ting says Votee is profitable “in the sense that our revenues exceed our costs,” funded largely through client contracts. The startup now boasts Allan Zeman, the Hong Kong tycoon responsible for growing the city’s Lan Kwai Fong nightlife district, as an advisor.

Votee’s ambitions are not limited to just Hong Kong. Ting says the startup is in “active discussions” with AI Singapore, the country’s AI research initiative, and plans to expand further into Southeast Asia. Beyond that, Ting wants to explore using AI to protect endangered languages in regions like East Asia, North America, and Africa.

Ting calls what English-language AI is doing to other languages a “typewriter moment”—a productivity gain so large that people abandon their own language to get it. “People will adopt English just because the typewriter’s productivity is so strong versus their own language,” he says.

Whether a 70-billion-parameter Cantonese model changes that remains to be seen. But Ting says he is still motivated by a drive to give other languages a fighting chance in a world dominated by English- and Mandarin Chinese-language AI.

“Every language that dies, you lose another way of seeing the world,” he says. “That could just be preserved in a museum where you can kind of see it. But we can also unlock a lot of new wisdom.”

Pak-Sun Ting will be speaking at the Fortune Leaders Forum, held in Macau on Sep. 8.

This story was originally featured on Fortune.