Now commissioning · TH-BKK-1
NVIDIA Blackwell Ultra, racked in Bangkok.
Forty B300 GPUs on Thai soil, under Thai law. Rent one for an hour, reserve a node for a year, or have us run a private AI platform end to end — without your data ever leaving the country.
Billed in USD or THB · No egress charge to Thai networks · Contracts under Thai law

40 × B300
Blackwell Ultra GPUs in Bangkok
10.5 TB
HBM3e across the fleet
720 PFLOPS
FP4 inference capacity
2–6 ms
Round trip from Thai networks
What we sell
The same forty GPUs, sold four different ways.
Raw GPU hours are the cheapest way in and the least valuable thing we do. Most customers start at the top of this list to prove the hardware, then move down it once they know what they actually need.

GPU Cloud
On-demand B300 instances, billed by the second.
Root on bare metal, CUDA and PyTorch images ready, no commitment. The fastest way to find out whether the hardware does what your benchmark needs.
From $8.90 per GPU-hour

Dedicated Nodes
A whole eight-GPU server, allocated to you alone.
No neighbours on the NVLink domain, no shared NVMe, no contention. One predictable line on the invoice every month.
From $30,900 per node / month

Private AI Cloud
A managed AI platform inside your own perimeter.
Single-tenant GPUs, private VPC, managed model serving, retrieval over your own systems, SSO and full audit. Built for banks, insurers and hospitals.
From $34,000 per month

Inference API
Open models on an OpenAI-compatible endpoint.
Point your existing SDK at a Bangkok endpoint. Prompts and completions stay in Thailand and are never used to train anything.
From $0.09 per million tokens

Sovereignty
“In the region” is not the same as “in the country”.
Every hyperscaler will sell you a Singapore or Jakarta region and call it local. Your board, your regulator and your customers are asking a narrower question — and it deserves a specific answer.
- One region, by design. Cross-border replication is not disabled by policy; it does not exist as a capability.
- Thai law, Thai courts. PDPA obligations sit with a Thai entity you can serve notice on.
- No foreign parent. No disclosure statute from another jurisdiction reaches the operator.
- Zero retention by default. Prompts are held in memory for the request and never written to disk.
Infrastructure
One node is a supercomputer. We have five.
Each node is an NVIDIA HGX B300 baseboard: eight GPUs on a fifth-generation NVLink switch, 2.1 TB of HBM3e in one coherent domain, 144 PFLOPS of FP4. A 671B mixture-of-experts model loads on a single chassis with room left for a long KV cache.
All five hang off the same Quantum-X800 InfiniBand fabric at 800 Gb/s per port, so a training job can span the fleet when you need forty GPUs acting as one.

Availability
Five nodes exist. We will tell you which are free.
Capacity is the one number infrastructure vendors are tempted to exaggerate. Publishing it keeps us honest and lets you plan against a real date rather than a sales assurance. Today 19 of 40 GPUs are unallocated.
- node-01Reserved
- node-02Reserved
- node-03Partial
- node-04Available
- node-05Available

Solutions
Most companies do not want GPUs. They want an answer.
If you have an AI team, take the compute and go. If you do not, these are the things we will build and then operate for you on the same fleet.
Enterprise RAG
Connectors for SharePoint, Google Drive, Confluence, S3, SQL and common Thai ERP and CRM systems.
Learn moreFrom $48,000 setup + from $7,500/month
Fine-tuning
Data preparation, LoRA or full fine-tune, held-out evaluation against your own rubric, then deployment onto your reserved capacity.
Learn moreFrom $14,500 per project
Managed LLM hosting
vLLM or TensorRT-LLM, autoscaling, request queueing, token quotas, canary rollouts, dashboards and 24×7 paging.
Learn moreFrom $5,900/month + compute
AI agent platform
Tool registry, sandboxed execution, human approval gates, per-action audit trail.
Learn moreFrom $4,200/month + usage
Vision AI
Defect detection, LPR, people counting, PPE and safety-zone compliance.
Learn moreFrom $11 per camera / month
Migration and landing zone
Network design, identity federation, data transfer, cutover plan and parallel run.
Learn moreFrom $18,000
How it starts
No discovery call. Bring a workload.
Typically two weeks from first email to running on real hardware.
Tell us the shape of the job
Model and parameter count, context length, tokens per second at peak, training or serving. A rough answer is enough to size it.
We return a sizing and a number
Within two business days: GPU count, expected throughput, monthly cost at your committed term, and the date capacity frees up. In writing.
Run a paid pilot
Two weeks on real hardware against your own evaluation set. Pilot spend is credited in full against a term contract.
Commit, or walk away
If the numbers hold, sign a 6, 12 or 36-month term. If they do not, you keep the results and owe nothing further.
Talk to us
19 of 40 GPUs are still unallocated.
Tell us the model, the context length and the tokens per second you need. You will get a sizing, a price and an availability date — in writing, within two business days.
Direct line sales@thaiaicloud.co.th · Replies in one business day, Thai or English.