Skip to main content

Templates

Pick a GPU, pick an image, go

Delete "set up the environment" from your workflow. Every template is a complete image with drivers, CUDA, frameworks and dependencies already installed, so the instance is usable the moment it boots — not a bare Ubuntu box waiting for you to pip install.

Steps from zero to running
3
Dependencies to install
0
GPU models to pair with
75
Per hour, from
$0.193

01 — Categories

Choose by what you are building

Same compute, different image, different production line. These six cover most of what runs here.

  • LLM inference serving

    vLLM, TGI, Ollama and SGLang preconfigured — boot the instance and you have an OpenAI-compatible HTTP endpoint. The fastest path from an open-weights model to your own private API.

    vLLM · TGI · Ollama · SGLang

  • Training and fine-tuning

    PyTorch, TensorFlow and JAX with matched CUDA and cuDNN, plus LLaMA Factory, Axolotl and Unsloth. LoRA, QLoRA and full-parameter runs work out of the box.

    PyTorch · TensorFlow · JAX · Axolotl

  • Image and video generation

    ComfyUI and AUTOMATIC1111 with the common nodes and samplers preinstalled, covering SDXL, FLUX and video diffusion models. The web UI is reachable in the browser on boot.

    ComfyUI · A1111 · FLUX · SDXL

  • Speech and audio

    The full Whisper family with GPU-accelerated transcription, including faster-whisper and WhisperX alignment. Batch audio processing and subtitle generation work immediately.

    Whisper · faster-whisper · WhisperX

  • Data processing and scientific computing

    The NVIDIA RAPIDS suite (cuDF, cuML, cuGraph) to move pandas-scale pipelines onto the GPU, plus Jupyter and the usual scientific stack.

    RAPIDS · cuDF · cuML · Jupyter

  • Clean systems and custom images

    Bare Ubuntu 22.04 or 24.04 CLI images carrying only drivers and the CUDA toolkit — the rest is yours. You can also point us at your own Docker image.

    Ubuntu · CUDA Toolkit · Custom image

02 — Deployment

Three steps from order to shell

No tickets, no quota approval, no waiting list.

  • 01

    Pick a machine

    Filter by VRAM, price, region, bandwidth and reliability against live inventory. The price you see is the price you pay, and it locks at order time.

  • 02

    Pick an image and disk size

    Choose a template or supply your own image. Size the disk for what you need — leave headroom for large model weights, since disk bills separately per GB-month.

  • 03

    Connect

    Once ready you get a direct SSH command, a Jupyter URL and the full port mapping table. You can also open a root shell in the browser or install a visual control panel with one click.

03 — Choosing a GPU

Which card will run your model

VRAM is the hard constraint. Below are half-precision minimums for common open-weight models and the cheapest card that clears each bar.

Which card will run your model
Model sizeWeightsMin VRAMSuggested GPU
DeepSeek V4Flash · 3-bit103GB110GBH200140GB$6.660/hr
DeepSeek V4Flash · 4-bit155GB162GBB200179GB$10.797/hr
DeepSeek V4Flash · 8-bit162GB169GBB200179GB$10.797/hr
DeepSeek V4Pro · Q2_K400GB420GBNeeds a multi-GPU node
Qwen3.827B · INT414GB20GBTesla V10032GB$0.188/hr
Qwen3.827B · INT828GB36GBQ RTX 800048GB$0.508/hr
Qwen3.827B · FP1656GB64GBA100 SXM480GB$1.088/hr
Qwen3.8Max · 量化400GB420GBNeeds a multi-GPU node

These are floors for loading FP16/BF16 weights; real deployments also need headroom for KV cache and batching. INT8 or INT4 quantisation lowers the bar substantially — a single 24GB RTX 4090 will serve a quantised 32B model.

04 — Ecosystem

The tools you already use, unchanged

Stock NVIDIA drivers and a standard CUDA environment. No patched runtime, no proprietary SDK to bind against. Code that runs on your workstation runs here.

  • PyTorch
  • TensorFlow
  • JAX
  • ComfyUI
  • Stable Diffusion
  • FLUX
  • vLLM
  • Ollama
  • Whisper
  • RAPIDS
  • Blender
  • Axolotl
  • Unsloth
  • LLaMA Factory

05 — FAQ

Images and environments

Can I use my own Docker image?

Yes. Supply the image reference when creating the instance; public registries and credentialed private registries are both supported, as are custom entrypoints and environment variables.

Which CUDA version ships in the images? Can I change it?

Each template carries the CUDA build that matches its framework version, with the PyTorch line tracking a recent stable branch. If your code pins a specific CUDA version, the reliable path is your own image — that puts the whole version matrix under your control.

Does data survive after an instance is destroyed?

No. Destroying an instance releases the disk and the data is unrecoverable. Copy results out first. Anything you need long-term should be synced to object storage or your own server.

Can I expose custom ports for my own service?

Yes. Every instance ships with a range of public direct ports, and the full mapping table is listed under advanced settings on the instance page. One caveat: port mappings are fixed when the container is created and cannot be added afterwards, so allocate enough ports at order time.

Is multi-GPU training supported?

Yes. You can filter the marketplace for multi-GPU nodes, up to 14 GPUs and 2152GB of combined VRAM in a single node. DeepSpeed, FSDP and Megatron work normally. For cross-node training, prefer nodes with high-speed interconnect.

Start building on NexGPU

Enterprise R&D team or solo developer — either way, your first job can be running in minutes.

Sign up to browse live network pricing. No payment method required.