Skip to main content

AI text generation

Turn an open model into your own API

Stop handing your data to a third party and stop paying per token. Rent a GPU, start vLLM, and you have an OpenAI-compatible endpoint running on a machine you control — concurrency, context length and quantisation all your call.

7B-class per hour, from
$0.540
70B-class per hour, from
$1.088
Max VRAM per node
2152GB

01 — When it applies

When self-hosting inference makes sense

Commercial APIs bill per token, which is excellent at low volume. Once you move to batch generation, long-context processing or a high-concurrency production service, the bill rises faster than most teams model — and the curve is not yours to control. The vendor can reprice, rate-limit or retire a model version at any time.

Self-hosted inference has a different cost structure entirely: you pay for GPU time, not token count. An RTX 4090 costs $0.540 per hour and will serve a quantised 7B model at meaningful concurrency. Once throughput is high enough, your cost per token falls to a fraction of an API's.

The sharper reason is the data boundary. Medical records, legal documents, internal knowledge bases, user conversation logs — once any of it leaves for an external endpoint it is out of your control. Self-hosting keeps it on the machine you rented, and destroying the instance erases it.

02 — Workloads

What it handles

One inference service; attach a different front end and it becomes a different product.

  • Private knowledge-base Q&A (RAG)

    Vectorise internal documents, product manuals and historical tickets, then wire them into the inference service to build an assistant that only answers questions about your business. Retrieval and generation both happen on the same machine — nothing leaves it.

  • Bulk content generation

    Product descriptions, summaries, translation, data labelling. Hundreds of thousands of rows is expensive per token and a fixed cost per GPU-hour. Run it, then destroy the instance.

  • Production chat services

    vLLM's continuous batching and PagedAttention let a single card carry real concurrency. Point your front end at it and both latency and throughput stay under your control.

  • Model evaluation and selection

    Boot several candidate models on the same machine in turn and compare them against your own test set. A few hours of rent buys an evidence-based decision.

03 — Stack

Inference engines, preinstalled

Three engines spanning maximum throughput to fastest time-to-running.

  • vLLM — throughput first

    PagedAttention and continuous batching set the open-source throughput benchmark. The right pick for concurrent production traffic and large offline generation runs. Serves an OpenAI-compatible /v1/chat/completions on boot.

    vLLM · OpenAI-compatible · Tensor parallel

  • Ollama — fastest to running

    One command pulls a model and handles quantisation and VRAM allocation automatically. Ideal for validating an idea, local development against a remote GPU, or simply wanting a working endpoint without tuning anything.

    Ollama · GGUF · Auto-quantisation

  • TGI and SGLang — targeted strengths

    TGI is more mature on streaming output and production observability; SGLang's RadixAttention wins clearly on multi-turn dialogue and structured prompts. Pick by the shape of your traffic.

    TGI · SGLang · RadixAttention

04 — Choosing a GPU

Which card for which model size

The hard constraint on inference is fitting weights plus KV cache in VRAM. Below are minimums for common deployment sizes and the cheapest card currently rentable at each.

DeploymentVRAM neededCheapest availableNotes
7B model, INT4 quantised10GBRTX 3060$0.100/hrPersonal assistants and small internal pilots. Quality loss from quantisation is modest; the best value entry point.
14B model, FP1632GBTesla V100$0.188/hrThe sweet spot for most production assistants — the best balance of quality against cost.
32B model, INT4 quantised24GBTesla V100$0.188/hrQuantisation squeezes 32B into 24GB — the highest quality reachable on a single GPU.
70B model, FP16141GBB200$10.797/hrWeights alone are 140GB, and a real deployment needs KV cache headroom on top, so this usually means a multi-GPU node. vLLM handles tensor parallelism itself; no code changes required.

"Cheapest available" is derived from live inventory — the lowest per-GPU rate among models that clear the VRAM bar — and moves with the market. Multi-GPU nodes rent whole.

05 — Getting started

Three steps to a running endpoint

  • 01

    Pick a card and reserve ports

    Order at the VRAM tier from the table above. Remember that port mappings are fixed at container creation, so allocate the port your inference service will expose.

  • 02

    Boot the vLLM image

    Choose the vLLM template and supply a model name; weights download and the server starts automatically. Pulling large weights takes a while, and only storage bills during that window.

  • 03

    Point your app at it

    The instance page gives you the mapped public address. Use it as your OpenAI base_url and the migration is a one-line change.

06 — FAQ

LLM inference deployment

What does hosting a 7B model cost per hour?

An INT4-quantised 7B model fits in 10GB, and the cheapest suitable card runs under $0.15 per hour. Even FP16 on an RTX 4090 is only $0.540 per hour. At eight hours a day, a month of inference lands in the tens of dollars.

vLLM or Ollama?

Serving real traffic and care about throughput and concurrency — vLLM. Just want a model running to see how it behaves — Ollama, which starts with one command at the cost of peak throughput. Both templates are preinstalled; switching means booting a new instance.

Do I have to upload the model weights myself?

No. Almost all open-weight models can be pulled straight from Hugging Face inside the instance over the host's public bandwidth. For your own fine-tuned private weights, sync them in with scp, rsync or object storage.

Can I run tensor-parallel multi-GPU inference?

Yes. Filter the marketplace for multi-GPU nodes and pass --tensor-parallel-size to vLLM; no model code changes are needed. Up to 14 GPUs and 2152GB of combined VRAM in one node.

Are model weights kept after the instance is destroyed?

No — destroying an instance releases the disk. If you reuse the same weights often, sync them to object storage and pull them back next time. You can also leave the instance stopped, but note that storage charges continue while it is stopped.

Start building on NexGPU

Enterprise R&D team or solo developer — either way, your first job can be running in minutes.

Sign up to browse live network pricing. No payment method required.