ILMU API: Malaysia's Sovereign AI Stack, Priced in Ringgit

ILMU API puts language, vision, speech, and embeddings on Malaysian soil, billed in ringgit. What YTL AI Labs' sovereign inference platform means for builders.

ILMU API: Malaysia's Sovereign AI Stack, Priced in Ringgit

What We're Seeing on the Ground

For the past two years, nearly every AI architecture review we've run for a Malaysian regulated client has stalled at the same wall. Not model quality. Not budget. A single question from the risk team: "Where does the data go?"

Until recently there were two honest answers, and both were compromises. Route inference to a frontier model offshore and spend months engineering around cross-border data transfer under the PDPA — or self-host open-weight models onshore and inherit a GPU cluster, an MLOps function, and a capacity-planning problem you never asked for.

ILMU API changes the default. YTL AI Labs has opened its sovereign inference platform to builders: language, vision, speech, and embedding models hosted on Malaysian infrastructure, callable from the OpenAI, Anthropic, or Gemini SDK you already use, and billed in ringgit (Source: ILMU API Docs — Overview).

Most people know ILMU as the consumer chat app. The API is the part that matters for enterprise builders — and this post is the ground-truth read, with the numbers.

What ILMU API Actually Is

image

Strip away the branding and ILMU API is a unified inference gateway in front of models running on Malaysian infrastructure. Three design decisions make it unusually easy to adopt:

One key, three SDK formats. The platform exposes OpenAI-compatible, Anthropic-compatible, and Gemini-compatible endpoints, all accepting the same sk- prefixed key (Source: ILMU API Docs — Overview):

SDK format

Base URL

OpenAI

https://api.ilmu.ai/v1

Anthropic

https://api.ilmu.ai/anthropic

Gemini

https://api.ilmu.ai/gemini

If your codebase already talks to any of the big three, migration is a base-URL change and a new key. No wrapper libraries, no SDK rewrite. For teams with existing LangChain, LlamaIndex, or plain-SDK integrations, that collapses the switching cost to a config value — which is exactly how a sovereign option becomes a practical one rather than a strategic aspiration.

Data residency by architecture, not by contract clause. Inference runs on Malaysian infrastructure, with low-latency access from Southeast Asia. Residency isn't a premium tier or a special deployment — it's the platform's default state.

Token-based billing with per-seat allocation. Organisations manage seats through a console, monitor spend in real time, and pay per token consumed — with no minimum fee, and failed requests never billed (Source: ILMU API Docs — Pricing).

The Stack Is Wider Than Chat

image

The part that changes real system design isn't the flagship chat model — it's that ILMU API ships the supporting cast most business workloads actually run on (Source: ILMU API Docs — Models):

Language. ilmu-v3.1 is the flagship: extended thinking, OpenAI-style function calling, a 200K-token context window, and native multilingual support. ilmu-mini-v3.3 is the high-volume workhorse for classification, summarisation, and simple Q&A — same 200K window at a fraction of the price. Two plan-exclusive agent models (nemo-super, ilmu-nemo-nano) push to 256K context with full tool-use, the nano variant carrying deeper Bahasa Malaysia training.

Vision with real OCR. ilmu-vision-v1.3 combines image analysis with high-accuracy OCR across printed, handwritten, and multilingual text in complex document layouts. Anyone who has processed Malaysian claims files, KYC documents, or invoices knows that "handwritten and multilingual" is not a nice-to-have — it's the job.

Speech built for how Malaysians actually talk. ilmu-asr-v4.2 transcribes English, Bahasa Malaysia, and Mandarin including natural code-switching — the Manglish reality every contact-centre transcript lives in. ilmu-tts-v2 generates natural voice output with strong Bahasa Malaysia support.

Embeddings and reranking. ilmu-embedding-v1 (32K input, 4,096 dimensions), bge-m3, and bge-reranker complete the retrieval loop — meaning an entire RAG pipeline, from embedding to rerank to generation, can now run without a single token leaving the country.

That last sentence is the headline for anyone building on sensitive corpora. The full retrieval-augmented stack — not just the chat endpoint — is onshore.

The Pricing Maths, in Ringgit

image

Here's what the data actually says. All rates below are pay-as-you-go list prices, in Malaysian Ringgit (Source: ILMU API Docs — Pricing):

Model

Role

Input

Output

ilmu-v3.1

Flagship reasoning

RM 4.00 / MTok

RM 16.00 / MTok

ilmu-vision-v1.3

Vision + OCR

RM 1.60 / MTok

RM 4.80 / MTok

ilmu-mini-v3.3

High-volume workhorse

RM 0.20 / MTok

RM 1.20 / MTok

ilmu-embedding-v1

Embeddings

RM 0.40 / MTok

bge-m3 / bge-reranker

Embeddings / rerank

RM 0.04 / MTok

ilmu-asr-v4.2

Speech-to-text

RM 0.0002 / second

ilmu-tts-v2

Text-to-speech

RM 0.08 / 1,000 chars

Run the numbers on real workloads and the scale becomes obvious. The docs' own worked example — a flagship call with 2,000 input and 500 output tokens — costs RM 0.016. At list rates, transcribing a full hour of call audio comes to about RM 0.72. Embedding a ten-million-token SOP library on bge-m3 is a one-off RM 0.40. These are rounding errors, not line items.

Three billing details matter more than the headline rates. The bill is in ringgit — your unit economics carry zero USD exposure, which any CFO who has watched an AI budget move with the exchange rate will appreciate. Reasoning tokens aren't billed separately — thinking is reported for audit but already counted within output tokens. And prepaid PAYG credit is org-wide with 12-month validity, so finance can treat it like any other utility.

We wrote recently about the four-layer total cost of ownership of AI products — build, run, govern, people. ILMU API attacks Layer 2, the run cost, at the pricing floor. The other three layers are still yours. More on that below.

What Residency Changes in the Compliance Conversation

image

Let's be precise, because this is where vendor marketing usually gets ahead of itself.

The amended PDPA tightened breach notification, introduced data protection officer requirements, and reworked the cross-border transfer regime (Source: Personal Data Protection Department). BNM's Risk Management in Technology (RMiT) expectations make every offshore inference endpoint a third-party technology-risk conversation for financial institutions (Source: Bank Negara Malaysia). And the National Guidelines on AI Governance and Ethics (AIGE) set the accountability bar for how AI systems are deployed (Source: National AI Office).

An onshore endpoint does not make your product compliant. You still own consent, purpose limitation, access control, audit trails, and human oversight. What residency does is remove the hardest structural question from the assessment — the one that turns a six-week review into a six-month one. When personal data never crosses the border for inference, the cross-border transfer analysis largely disappears, and the outsourcing-risk conversation gets dramatically shorter.

And ILMU is not arriving unproven in regulated territory. Ryt Bank — the BNM-licensed digital bank from YTL and Sea — launched in 2025 with ILMU as the model behind its customer-facing AI (Source: Ryt Bank). A homegrown model already operates inside the perimeter of Malaysian banking supervision. That precedent is worth more to a compliance committee than any benchmark chart.

For context on why the stakes are rising: Malaysia drew RM87.4 billion in AI-driven digital investments in 2025 under the AI Nation 2030 agenda (Source: The Edge Malaysia), and MDEC's MDAG-AI grant co-funds up to 70% of qualifying AI builds, capped at RM2 million (Source: MDEC). The capital and the policy support are here. What was missing was a sovereign inference layer priced for production. That gap just closed.

The Symprio Approach: Sovereign by Design, Routed by Data Class

Symprio builds AI products for Malaysian regulated industries — and advises on the architecture underneath them. ILMU API slots directly into how we already work. Four principles guide it.

Route by data class, not by hype. Our composable enterprise architecture → treats model choice as a routing table: personal and regulated data to onshore endpoints, commodity tasks to the cheapest capable model, frontier reasoning only where the workload earns it. ILMU API gives that table a sovereign row with real production pricing.

Prove it with evals, in Bahasa and English. We build domain evaluation sets — code-switched transcripts, scanned Malaysian documents, regulator-grade edge cases — before any model touches production traffic. If ilmu-vision-v1.3 reads your handwritten claim forms better than your incumbent OCR, you'll know from evidence, not a brochure.

Co-build so the capability stays. Our adopt-and-build model pairs Symprio engineers with your team, transfers the architecture, and certifies your people to run it. A sovereign stack you can't operate yourself is just a different kind of dependency.

Govern from sprint one. Sovereign-cloud deployments → aligned to BNM expectations, PDPA, and AIGE from the first sprint — because retrofitted compliance is the most expensive kind. An onshore model makes that work easier; it doesn't make it optional.

This is precisely the work of our agentic AI products and platforms practice →: production agents on architecture that survives a regulator's questions.

What This Looks Like in Practice

Within 90 days of engagement, teams typically see deployments such as:

  • A claims intake co-pilot using ilmu-vision-v1.3 to OCR handwritten, multilingual claim documents — extraction for sen per file, with no image leaving Malaysia.

  • A contact-centre triage agent transcribing code-switched EN/BM/Mandarin calls on ilmu-asr-v4.2 (about RM0.72 per hour of audio at list rates) and classifying intent on ilmu-mini-v3.3.

  • An internal knowledge assistant over SOPs and policy manuals — bge-m3 embeddings, reranked retrieval, generation on the mini model — an entirely onshore RAG loop.

  • An AML investigation acceleration agent that reads a full case file in one pass inside ilmu-v3.1's 200K context window, with every action auditable.

  • An LHDN MyInvois e-invoicing middleware using vision extraction to bridge supplier documents into compliant submissions.

These are not moonshots. They're owned assets with run costs you can now quote in sen — on infrastructure your compliance team can actually approve.

Working With Us

We work with engineering and product teams in three modes:

  • Architecture review — a half-day deep-dive on your current stack, including where a sovereign inference layer fits, with a written recommendation within a week.

  • Embedded build — one or two Symprio engineers join your team for 8–12 weeks to co-build a specific AI-enabled product to production.

  • Platform partnership — longer-term engagement to ship a portfolio of products on a routing architecture you own.

👉 Explore our Agentic AI products & platforms → 👉 Book a 30-minute discovery call → — no slide deck, just whiteboard thinking 👉 Read related: Your AI Pilot Was the Cheap Part — The Real TCO of AI Products →

image

FAQ: ILMU API and Sovereign AI in Malaysia

What is ILMU API?

ILMU API is a sovereign AI inference platform built by YTL AI Labs, providing language, vision, speech, and embedding models hosted on Malaysian infrastructure. It exposes OpenAI-, Anthropic-, and Gemini-compatible endpoints under one API key, with token-based billing in ringgit.

How much does ILMU API cost?

Pay-as-you-go rates start at RM0.20 per million input tokens for the mini model and RM4.00 for the flagship, with embeddings from RM0.04 per million tokens, transcription at RM0.0002 per second of audio, and text-to-speech at RM0.08 per 1,000 characters. Failed requests are never billed.

Does ILMU API make my product PDPA-compliant?

Not by itself. Malaysian data residency removes the cross-border transfer question and shortens BNM-style outsourcing risk assessments, but consent, purpose limitation, access controls, audit trails, and AIGE-aligned governance remain the deployer's responsibility.

Can I use my existing OpenAI or Anthropic code with ILMU API?

Yes. ILMU API exposes compatible endpoints for the OpenAI, Anthropic, and Gemini SDK formats — you point your existing SDK at the ILMU base URL with an ILMU API key, with no wrapper libraries or SDK changes required.

Who built ILMU?

ILMU is built by YTL AI Labs in Malaysia. The model family also powers Ryt Bank, the BNM-licensed digital bank launched by YTL and Sea in 2025 — an early proof point of the stack operating in a regulated environment.

Sources & Further Reading

  1. ILMU API Documentation — Overview

  2. ILMU API Documentation — Models

  3. ILMU API Documentation — Pricing

  4. ILMU — Official Site

  5. YTL Corporation

  6. Ryt Bank

  7. Personal Data Protection Department, Malaysia

  8. Bank Negara Malaysia — Risk Management in Technology (RMiT)

  9. Malaysia National AI Office (NAIO) & AIGE Guidelines

  10. MDEC — Malaysia Digital Acceleration Grant: Artificial Intelligence (MDAG-AI)

  11. The Edge Malaysia — AI Nation 2030 and Malaysia's Next Phase of Growth

#Symprio #BuildNotBuy #EnterpriseAI #SovereignCloud #AgenticAI #PDPA #Malaysia


Symprio builds AI products on architecture that survives a regulator's questions — sovereign by design, owned by you. Find us at symprio.com.