Products
Compute, your way
Run the metal directly or go serverless over an API. Both are production-grade, security-first, and 10% below market.
GPU rental
Dedicated NVIDIA H100 & B300
Full-tenant bare metal with no virtualization overhead and no noisy neighbors. Built for training, fine-tuning, and heavy batch inference.
NVIDIA H100 SXM 80GB
Available today- 80GB HBM3 · 3.35 TB/s memory bandwidth
- NVLink + 400G InfiniBand inter-node fabric
- Single-node (1×) to multi-rack (256×) clusters
- Pre-warmed ML images: PyTorch, cuDNN, NCCL
- 4 pricing tiers, down to $2.03/GPU-hr on 12-month
NVIDIA B300 · Blackwell Ultra
Pre-sale — lock your rate- 288GB HBM3e · next-gen Blackwell Ultra
- Substantially higher throughput per dollar vs. Hopper
- Reserve capacity ahead of public availability
- Prepay to lock both the slot and the hourly rate
- Pre-sale at $7.85/GPU-hr (market rate)
Bare-metal tenancy
Dedicated nodes, isolated from other tenants end to end.
High-speed fabric
400G InfiniBand for near-linear multi-node scaling.
US regions
Low-latency capacity in multiple US data centers.
Your stack, your way
Bring your own images, or start from ours in minutes.
LLM token API
Serverless inference, OpenAI-compatible
Ship LLM features without managing a single server. Drop in our endpoint, keep your code unchanged, and pay per token.
- Drop-in compatible — same request/response shape as OpenAI's chat completions.
- Per-token billing — no minimums, no idle charges, scale to zero.
- Sub-second TTFT — optimized serving for low latency at production traffic.
- Kimi K3, DeepSeek V4 Pro & V4 Flash — served on our own GPUs.
- Zero data retention — your prompts and completions are never stored or trained on.
curl https://api.vetulink.com/v1/chat/completions \
-H "Authorization: Bearer $VETU_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "vetu/kimi-k3",
"messages": [{"role": "user",
"content": "Explain GPU memory bandwidth"}]
}'
# 200 OK · OpenAI-compatible · 35ms TTFT
Which is right for you?
Match the product to the workload
| Dedicated GPU rental | Token API | |
|---|---|---|
| Best for | Training & fine-tuning | Production inference |
| Control | Full stack access | Managed for you |
| Billing | Per GPU-hour | Per token |
| Setup time | Under an hour | Minutes (API key) |
| Idle cost | Paid while reserved | $0 (scale to zero) |
| Hardware | H100 · B300 | H100-backed fleet |
Not sure what you need?
Tell us about your workload and we'll recommend the right setup — and the right price.