Products

Compute, your way

Run the metal directly or go serverless over an API. Both are production-grade, security-first, and 10% below market.

GPU rental

Dedicated NVIDIA H100 & B300

Full-tenant bare metal with no virtualization overhead and no noisy neighbors. Built for training, fine-tuning, and heavy batch inference.

NVIDIA H100 SXM 80GB

Available today
  • 80GB HBM3 · 3.35 TB/s memory bandwidth
  • NVLink + 400G InfiniBand inter-node fabric
  • Single-node (1×) to multi-rack (256×) clusters
  • Pre-warmed ML images: PyTorch, cuDNN, NCCL
  • 4 pricing tiers, down to $2.03/GPU-hr on 12-month

NVIDIA B300 · Blackwell Ultra

Pre-sale — lock your rate
  • 288GB HBM3e · next-gen Blackwell Ultra
  • Substantially higher throughput per dollar vs. Hopper
  • Reserve capacity ahead of public availability
  • Prepay to lock both the slot and the hourly rate
  • Pre-sale at $7.85/GPU-hr (market rate)

Bare-metal tenancy

Dedicated nodes, isolated from other tenants end to end.

High-speed fabric

400G InfiniBand for near-linear multi-node scaling.

US regions

Low-latency capacity in multiple US data centers.

Your stack, your way

Bring your own images, or start from ours in minutes.

LLM token API

Serverless inference, OpenAI-compatible

Ship LLM features without managing a single server. Drop in our endpoint, keep your code unchanged, and pay per token.

  • Drop-in compatible — same request/response shape as OpenAI's chat completions.
  • Per-token billing — no minimums, no idle charges, scale to zero.
  • Sub-second TTFT — optimized serving for low latency at production traffic.
  • Kimi K3, DeepSeek V4 Pro & V4 Flash — served on our own GPUs.
  • Zero data retention — your prompts and completions are never stored or trained on.
curl https://api.vetulink.com/v1/chat/completions \
  -H "Authorization: Bearer $VETU_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "vetu/kimi-k3",
    "messages": [{"role": "user",
                 "content": "Explain GPU memory bandwidth"}]
  }'

# 200 OK · OpenAI-compatible · 35ms TTFT
Which is right for you?

Match the product to the workload

Dedicated GPU rental Token API
Best forTraining & fine-tuningProduction inference
ControlFull stack accessManaged for you
BillingPer GPU-hourPer token
Setup timeUnder an hourMinutes (API key)
Idle costPaid while reserved$0 (scale to zero)
HardwareH100 · B300H100-backed fleet

Not sure what you need?

Tell us about your workload and we'll recommend the right setup — and the right price.