AI text generation
Turn an open model into your own API
Stop handing your data to a third party and stop paying per token. Rent a GPU, start vLLM, and you have an OpenAI-compatible endpoint running on a machine you control — concurrency, context length and quantisation all your call.
- 7B-class per hour, from
- $0.540
- 70B-class per hour, from
- $1.088
- Max VRAM per node
- 2152GB
01 — When it applies
When self-hosting inference makes sense
Commercial APIs bill per token, which is excellent at low volume. Once you move to batch generation, long-context processing or a high-concurrency production service, the bill rises faster than most teams model — and the curve is not yours to control. The vendor can reprice, rate-limit or retire a model version at any time.
Self-hosted inference has a different cost structure entirely: you pay for GPU time, not token count. An RTX 4090 costs $0.540 per hour and will serve a quantised 7B model at meaningful concurrency. Once throughput is high enough, your cost per token falls to a fraction of an API's.
The sharper reason is the data boundary. Medical records, legal documents, internal knowledge bases, user conversation logs — once any of it leaves for an external endpoint it is out of your control. Self-hosting keeps it on the machine you rented, and destroying the instance erases it.
02 — Workloads
What it handles
One inference service; attach a different front end and it becomes a different product.
Private knowledge-base Q&A (RAG)
Vectorise internal documents, product manuals and historical tickets, then wire them into the inference service to build an assistant that only answers questions about your business. Retrieval and generation both happen on the same machine — nothing leaves it.
Bulk content generation
Product descriptions, summaries, translation, data labelling. Hundreds of thousands of rows is expensive per token and a fixed cost per GPU-hour. Run it, then destroy the instance.
Production chat services
vLLM's continuous batching and PagedAttention let a single card carry real concurrency. Point your front end at it and both latency and throughput stay under your control.
Model evaluation and selection
Boot several candidate models on the same machine in turn and compare them against your own test set. A few hours of rent buys an evidence-based decision.
03 — Stack
Inference engines, preinstalled
Three engines spanning maximum throughput to fastest time-to-running.
vLLM — throughput first
PagedAttention and continuous batching set the open-source throughput benchmark. The right pick for concurrent production traffic and large offline generation runs. Serves an OpenAI-compatible /v1/chat/completions on boot.
vLLM · OpenAI-compatible · Tensor parallel
Ollama — fastest to running
One command pulls a model and handles quantisation and VRAM allocation automatically. Ideal for validating an idea, local development against a remote GPU, or simply wanting a working endpoint without tuning anything.
Ollama · GGUF · Auto-quantisation
TGI and SGLang — targeted strengths
TGI is more mature on streaming output and production observability; SGLang's RadixAttention wins clearly on multi-turn dialogue and structured prompts. Pick by the shape of your traffic.
TGI · SGLang · RadixAttention
04 — Choosing a GPU
Which card for which model size
The hard constraint on inference is fitting weights plus KV cache in VRAM. Below are minimums for common deployment sizes and the cheapest card currently rentable at each.
| Deployment | VRAM needed | Cheapest available | Notes |
|---|---|---|---|
| 7B model, INT4 quantised | 10GB | RTX 3060$0.100/hr | Personal assistants and small internal pilots. Quality loss from quantisation is modest; the best value entry point. |
| 14B model, FP16 | 32GB | Tesla V100$0.188/hr | The sweet spot for most production assistants — the best balance of quality against cost. |
| 32B model, INT4 quantised | 24GB | Tesla V100$0.188/hr | Quantisation squeezes 32B into 24GB — the highest quality reachable on a single GPU. |
| 70B model, FP16 | 141GB | B200$10.797/hr | Weights alone are 140GB, and a real deployment needs KV cache headroom on top, so this usually means a multi-GPU node. vLLM handles tensor parallelism itself; no code changes required. |
"Cheapest available" is derived from live inventory — the lowest per-GPU rate among models that clear the VRAM bar — and moves with the market. Multi-GPU nodes rent whole.
05 — Getting started
Three steps to a running endpoint
01
Pick a card and reserve ports
Order at the VRAM tier from the table above. Remember that port mappings are fixed at container creation, so allocate the port your inference service will expose.
02
Boot the vLLM image
Choose the vLLM template and supply a model name; weights download and the server starts automatically. Pulling large weights takes a while, and only storage bills during that window.
03
Point your app at it
The instance page gives you the mapped public address. Use it as your OpenAI base_url and the migration is a one-line change.
06 — FAQ
LLM inference deployment
What does hosting a 7B model cost per hour?
vLLM or Ollama?
Do I have to upload the model weights myself?
Can I run tensor-parallel multi-GPU inference?
Are model weights kept after the instance is destroyed?
Related solutions
Other ways to use it
Same compute network — swap the image and it becomes a different production line.
AI Agents
Deploy and scale agents on LangChain, CrewAI
Private LLM Deployment
Open weights, running on your own machine
AI Fine-tuning
LoRA, QLoRA and full fine-tuning, on demand
AI Image & Video
Stable Diffusion, FLUX and ComfyUI, ready to run
AI/ML Frameworks
Native PyTorch, TensorFlow and JAX
Audio to Text
GPU-accelerated Whisper transcription
Start building on NexGPU
Enterprise R&D team or solo developer — either way, your first job can be running in minutes.
Sign up to browse live network pricing. No payment method required.
