AI/ML frameworks
Skip the three days of setup
The driver does not match cuDNN, cuDNN does not match PyTorch, and PyTorch does not match the paper you are trying to reproduce. Everyone in deep learning has been through that loop. A prebuilt image deletes it.
- Single-GPU training, from
- $0.540
- Max GPUs per node
- 14
- Dependencies to configure
- 0
01 — When it applies
Why environment setup eats so much time
Deep learning has a far longer dependency chain than ordinary software: GPU driver, CUDA runtime, cuDNN, NCCL, the framework itself, plus a set of compiled extensions like xformers, flash-attention and apex. When any two links mismatch, the error usually surfaces somewhere entirely unrelated — which is what makes it so hard to diagnose.
Reproducing someone else's work is worse. A two-year-old paper used the PyTorch and CUDA of its day; running it on a current stack means a chain of API changes and subtly different operator behaviour. What you need there is not the newest environment but the ability to pin an exact one.
Prebuilt images solve the first problem: stable, tested version combinations for the common frameworks, ready to train on boot. Custom images solve the second: point us at any Docker image and reproduce the author's environment exactly, with the whole version matrix under your control.
02 — Workloads
Shapes of training work
From getting one GPU working to scaling out, from reproducing papers to training from scratch.
Single-GPU experiments and tuning
Changing architecture, sweeping hyperparameters, trying loss functions — work that needs fast iteration rather than peak compute. A $0.540/hr 4090 is plenty; destroy it when the sweep ends.
Multi-GPU data parallelism
When the dataset is large and one epoch takes too long, DDP or FSDP splits the batch across cards. Filter the marketplace for multi-GPU nodes, up to 14 per node.
Paper reproduction
Use a custom image to restore the author's environment precisely and avoid result drift from version skew. This is a quiet advantage of renting over a fixed local install.
Computer vision and classical models
Detection, segmentation, super-resolution, point clouds — not every GPU job is an LLM. These have modest VRAM needs, which is exactly where consumer cards win hardest on value.
03 — Stack
All three frameworks, preinstalled
Stock NVIDIA drivers and CUDA, no patched runtime — it behaves the way it does on your workstation.
PyTorch
With matched CUDA and cuDNN, plus torchvision, torchaudio and the usual compiled extensions. Launch distributed jobs with torchrun; DDP and FSDP work out of the box.
PyTorch · torchrun · DDP/FSDP
TensorFlow / Keras
GPU build installed with device visibility verified, including Keras 3. The right pick for existing TF codebases and SavedModel deployment paths.
TensorFlow · Keras 3 · SavedModel
JAX and distributed extras
JAX on the CUDA backend with Flax and Optax. DeepSpeed and Accelerate are in the image too, so large-model training needs no extra installation.
JAX · Flax · DeepSpeed · Accelerate
04 — Choosing a GPU
Sizing a card for training
Training is hungrier than inference — beyond weights you hold gradients, optimiser state and activations. As a rule of thumb it needs three to four times the inference footprint.
| Training scenario | VRAM needed | Cheapest available | Notes |
|---|---|---|---|
| CV models and small networks | 12GB | RTX 3060$0.100/hr | Detection, segmentation and classification have modest VRAM needs; consumer cards are entirely sufficient. |
| Mid-size models, large batches | 24GB | Tesla V100$0.188/hr | For jobs that need a larger batch size to converge stably. 24GB is the value sweet spot. |
| Large-model single-GPU training | 48GB | Q RTX 8000$0.508/hr | The 48GB tier, which with gradient checkpointing trains a substantial model. |
| Distributed training nodes | 80GB | A100 SXM4$1.088/hr | The 80GB tier at minimum. For multi-GPU, prefer nodes with high-speed interconnect — communication overhead varies enormously. |
"Cheapest available" is derived from live inventory — the lowest per-GPU rate among models that clear the VRAM bar — and moves with the market. Multi-GPU nodes rent whole.
05 — Getting started
Your first training run
01
Pick a framework image
PyTorch, TensorFlow or JAX. If your code pins a specific version combination, supplying your own Docker image is the safer route.
02
Get code and data across
git clone from the web terminal, or scp from your machine. For large datasets, stage them in object storage and pull from inside the instance — the host's bandwidth is far faster.
03
Train and checkpoint
Configure periodic checkpointing in your script, download the weights when it finishes, then destroy the instance. Destroying releases the disk; there is no recycle bin.
06 — FAQ
Deep learning training
How much VRAM does training need, and how do I estimate it?
How much faster is multi-GPU?
Can I pin the CUDA version?
How do I keep a long training run from being interrupted?
What is the fastest way to upload a dataset?
Related solutions
Other ways to use it
Same compute network — swap the image and it becomes a different production line.
AI Agents
Deploy and scale agents on LangChain, CrewAI
Private LLM Deployment
Open weights, running on your own machine
AI Fine-tuning
LoRA, QLoRA and full fine-tuning, on demand
AI Image & Video
Stable Diffusion, FLUX and ComfyUI, ready to run
AI Text Generation
vLLM, TGI and Ollama, live in minutes
Audio to Text
GPU-accelerated Whisper transcription
Start building on NexGPU
Enterprise R&D team or solo developer — either way, your first job can be running in minutes.
Sign up to browse live network pricing. No payment method required.
