Home/Products
Four ways to buy the same fleet
From one GPU-hour to a managed AI platform.
Every product below runs on the same five HGX B300 nodes in Bangkok. What changes is how much of the operating burden you keep, and how predictable you want the bill to be.
Product 01
GPU Cloud
One GPU, one minute, one command.
Billed per second from the moment the instance is reachable. GPUs are passed through, not time-sliced — no MIG partitioning unless you ask for it, no shared NVMe, no hypervisor tax on the memory bandwidth you paid for.
- Per-second billing with no minimum and no charge while provisioning
- Root on bare metal, or any OCI image from your own registry
- Free egress to BKNIX peers and Thai ISPs, 10 TB international included
- Persistent volumes that outlive the instance, so checkpoints survive an overnight shutdown

| Instance | GPU | GPU memory | vCPU | RAM | NVMe | Network | USD / hour |
|---|---|---|---|---|---|---|---|
| tac.b300.1x | 1 × B300 | 288 GB | 28 | 256 GB | 3.8 TB | 100 Gb/s | $8.90 |
| tac.b300.2x | 2 × B300 | 576 GB | 56 | 512 GB | 7.7 TB | 200 Gb/s | $17.80 |
| tac.b300.4x | 4 × B300 | 1.15 TB | 112 | 1 TB | 15.4 TB | 400 Gb/s | $35.60 |
| tac.b300.8x | 8 × B300 | 2.1 TB | 224 | 2 TB | 30.7 TB | 800 Gb/s | $69.90 |

Product 02
Dedicated Nodes
A whole HGX B300. Nobody else on it.
Eight Blackwell Ultra GPUs, 2.1 TB of HBM3e and 224 CPU cores allocated to one tenant for the length of a term. The NVLink domain stays whole, so the largest open models load without sharding across servers and no inter-node hop lands in the token path.
- Sole tenancy on GPU, CPU, memory bus, NVMe and NICs
- One predictable line on the invoice, agreed before the term starts
- A named machine with a serial number and a rack position, which is what auditors accept
- 48-hour provisioning from signature, with a spare node held in region
| Term | Per node / month | Per GPU-hour |
|---|---|---|
| Monthly | $46,900 | $8.03 |
| 6 months | $40,900 | $7.00 |
| 12 months | $35,900 | $6.15 |
| 36 months | $30,900 | $5.29 |
| Node specification | |
|---|---|
| GPU | 8 × NVIDIA B300 (Blackwell Ultra) |
| GPU memory | 2.1 TB HBM3e aggregate |
| NVLink | 14.4 TB/s all-to-all, 5th generation |
| Inference | 144 PFLOPS FP4 |
| Training | 72 PFLOPS FP8 |
| CPU | 2 × 112-core x86, 224 cores total |
Full specification, fabric and facility detail on theinfrastructure page.
Product 03
Private AI Cloud
Your own AI platform, run by us, inside your perimeter.
For institutions that cannot put their data on someone else's inference API and do not want to build a GPU operations team to avoid it. Single tenant, in Bangkok, with a Thai entity on the contract.
Most “private” AI offerings are a logical partition on shared silicon. This is not that — single tenancy here means the hardware, not the namespace.

Foundation
A first production LLM workload behind your own perimeter.
$34,000 / month
Onboarding from $25,000 · SLA 99.5%
- 4 dedicated B300, single tenant
- Private VPC, no shared data plane
- One managed open-weight model, served on vLLM or TensorRT-LLM
- S3-compatible object storage, 20 TB
- SSO via SAML or OIDC
- Business-hours Thai support, 4 h response
Enterprise
Most chosenCompany-wide AI with retrieval over internal systems.
$72,000 / month
Onboarding from $45,000 · SLA 99.9%
- 8+ dedicated B300, single tenant
- Multi-model serving with routing and quotas
- Managed RAG platform with connectors
- Role-based access control, per-department budgets
- Full prompt and response audit log, exportable
- 24×7 Thai support, 1 h response, named engineer
Sovereign
Regulated institutions and public sector.
By negotiation
Scoped per institution · SLA 99.95%
- Dedicated cage or customer-owned rack
- Air-gapped deployment option, no egress path
- Hardware and staff vetted to your standard
- Evidence pack for BOT, OIC and PDPC examination
- Source-available control plane for inspection
- Exit plan with full data and weight handover

Product 04
Inference API
Change one base URL. Keep your data in the country.
Open-weight models served on the same fleet behind an OpenAI-compatible endpoint that terminates in Bangkok. Existing SDKs, existing libraries, existing observability — the only thing that moves is where the request lands.
- Zero retention by default — prompts live in memory for the request and are never written to disk
- Never used for training, contractually rather than as a console setting
- Thai treated as a first-class case, with published token ratios so you can predict cost in Thai
- Free tier of 5M tokens a month, no card required
| Model class | Size | Input / 1M | Output / 1M | Typical use |
|---|---|---|---|---|
| Compact | ≤ 9B | $0.09 | $0.25 | Classification, extraction, routing |
| Standard | 30 – 80B | $0.38 | $0.95 | General assistants, summarisation |
| Frontier MoE | 200B+ MoE | $0.65 | $2.20 | Hard reasoning, long context |
| Reasoning | Long CoT | $0.75 | $2.90 | Agents, multi-step planning |
| Vision-language | Multimodal | $0.45 | $1.20 | Document and image understanding |
Cached input −90%
Repeated prefixes billed at one tenth of the input rate.
Batch API −50%
Asynchronous jobs returned within 24 hours.
Provisioned throughput Flat
Reserve fixed tokens per second on dedicated GPUs. From $6,400/month.
Free tier $0
5M tokens per month for evaluation, no card required.
Committed terms
Utilisation is our problem. Discount is how we buy it from you.
A committed term lets us plan power, staffing and the next hardware order. That is worth real money to us, so we hand most of it back.
| Term | USD / GPU-hour | Discount | 8-GPU node / month | Included |
|---|---|---|---|---|
| On-demand | $8.90 | — | $51,976 | No commitment, per second billing |
| 3 months | $7.60 | −15% | $44,384 | Capacity held, cancel at term |
| 6 months | $6.90 | −22% | $40,296 | Priority scheduling |
| 12 months | $6.20 | −30% | $36,208 | Named GPUs, quarterly true-up |
| 36 months | $5.30 | −40% | $30,952 | Hardware refresh clause included |
Talk to us
19 of 40 GPUs are still unallocated.
Tell us the model, the context length and the tokens per second you need. You will get a sizing, a price and an availability date — in writing, within two business days.
Direct line sales@thaiaicloud.co.th · Replies in one business day, Thai or English.