TACThai AI Cloud

Home/Products

Four ways to buy the same fleet

From one GPU-hour to a managed AI platform.

Every product below runs on the same five HGX B300 nodes in Bangkok. What changes is how much of the operating burden you keep, and how predictable you want the bill to be.

Product 01

GPU Cloud

One GPU, one minute, one command.

Billed per second from the moment the instance is reachable. GPUs are passed through, not time-sliced — no MIG partitioning unless you ask for it, no shared NVMe, no hypervisor tax on the memory bandwidth you paid for.

  • Per-second billing with no minimum and no charge while provisioning
  • Root on bare metal, or any OCI image from your own registry
  • Free egress to BKNIX peers and Thai ISPs, 10 TB international included
  • Persistent volumes that outlive the instance, so checkpoints survive an overnight shutdown
InstanceGPUGPU memoryvCPURAMNVMeNetworkUSD / hour
tac.b300.1x1 × B300288 GB28256 GB3.8 TB100 Gb/s$8.90
tac.b300.2x2 × B300576 GB56512 GB7.7 TB200 Gb/s$17.80
tac.b300.4x4 × B3001.15 TB1121 TB15.4 TB400 Gb/s$35.60
tac.b300.8x8 × B3002.1 TB2242 TB30.7 TB800 Gb/s$69.90

Product 02

Dedicated Nodes

A whole HGX B300. Nobody else on it.

Eight Blackwell Ultra GPUs, 2.1 TB of HBM3e and 224 CPU cores allocated to one tenant for the length of a term. The NVLink domain stays whole, so the largest open models load without sharding across servers and no inter-node hop lands in the token path.

  • Sole tenancy on GPU, CPU, memory bus, NVMe and NICs
  • One predictable line on the invoice, agreed before the term starts
  • A named machine with a serial number and a rack position, which is what auditors accept
  • 48-hour provisioning from signature, with a spare node held in region
TermPer node / monthPer GPU-hour
Monthly$46,900$8.03
6 months$40,900$7.00
12 months$35,900$6.15
36 months$30,900$5.29
Node specification
GPU8 × NVIDIA B300 (Blackwell Ultra)
GPU memory2.1 TB HBM3e aggregate
NVLink14.4 TB/s all-to-all, 5th generation
Inference144 PFLOPS FP4
Training72 PFLOPS FP8
CPU2 × 112-core x86, 224 cores total

Full specification, fabric and facility detail on theinfrastructure page.

Product 03

Private AI Cloud

Your own AI platform, run by us, inside your perimeter.

For institutions that cannot put their data on someone else's inference API and do not want to build a GPU operations team to avoid it. Single tenant, in Bangkok, with a Thai entity on the contract.

Most “private” AI offerings are a logical partition on shared silicon. This is not that — single tenancy here means the hardware, not the namespace.

Foundation

A first production LLM workload behind your own perimeter.

$34,000 / month

Onboarding from $25,000 · SLA 99.5%

  • 4 dedicated B300, single tenant
  • Private VPC, no shared data plane
  • One managed open-weight model, served on vLLM or TensorRT-LLM
  • S3-compatible object storage, 20 TB
  • SSO via SAML or OIDC
  • Business-hours Thai support, 4 h response

Discuss this tier

Enterprise

Most chosen

Company-wide AI with retrieval over internal systems.

$72,000 / month

Onboarding from $45,000 · SLA 99.9%

  • 8+ dedicated B300, single tenant
  • Multi-model serving with routing and quotas
  • Managed RAG platform with connectors
  • Role-based access control, per-department budgets
  • Full prompt and response audit log, exportable
  • 24×7 Thai support, 1 h response, named engineer

Discuss this tier

Sovereign

Regulated institutions and public sector.

By negotiation

Scoped per institution · SLA 99.95%

  • Dedicated cage or customer-owned rack
  • Air-gapped deployment option, no egress path
  • Hardware and staff vetted to your standard
  • Evidence pack for BOT, OIC and PDPC examination
  • Source-available control plane for inspection
  • Exit plan with full data and weight handover

Discuss this tier

Product 04

Inference API

Change one base URL. Keep your data in the country.

Open-weight models served on the same fleet behind an OpenAI-compatible endpoint that terminates in Bangkok. Existing SDKs, existing libraries, existing observability — the only thing that moves is where the request lands.

  • Zero retention by default — prompts live in memory for the request and are never written to disk
  • Never used for training, contractually rather than as a console setting
  • Thai treated as a first-class case, with published token ratios so you can predict cost in Thai
  • Free tier of 5M tokens a month, no card required
Model classSizeInput / 1MOutput / 1MTypical use
Compact≤ 9B$0.09$0.25Classification, extraction, routing
Standard30 – 80B$0.38$0.95General assistants, summarisation
Frontier MoE200B+ MoE$0.65$2.20Hard reasoning, long context
ReasoningLong CoT$0.75$2.90Agents, multi-step planning
Vision-languageMultimodal$0.45$1.20Document and image understanding

Cached input −90%

Repeated prefixes billed at one tenth of the input rate.

Batch API −50%

Asynchronous jobs returned within 24 hours.

Provisioned throughput Flat

Reserve fixed tokens per second on dedicated GPUs. From $6,400/month.

Free tier $0

5M tokens per month for evaluation, no card required.

Committed terms

Utilisation is our problem. Discount is how we buy it from you.

A committed term lets us plan power, staffing and the next hardware order. That is worth real money to us, so we hand most of it back.

TermUSD / GPU-hourDiscount8-GPU node / monthIncluded
On-demand$8.90$51,976No commitment, per second billing
3 months$7.60−15%$44,384Capacity held, cancel at term
6 months$6.90−22%$40,296Priority scheduling
12 months$6.20−30%$36,208Named GPUs, quarterly true-up
36 months$5.30−40%$30,952Hardware refresh clause included