Skip to main content

AI/ML frameworks

Skip the three days of setup

The driver does not match cuDNN, cuDNN does not match PyTorch, and PyTorch does not match the paper you are trying to reproduce. Everyone in deep learning has been through that loop. A prebuilt image deletes it.

Single-GPU training, from
$0.540
Max GPUs per node
14
Dependencies to configure
0

01 — When it applies

Why environment setup eats so much time

Deep learning has a far longer dependency chain than ordinary software: GPU driver, CUDA runtime, cuDNN, NCCL, the framework itself, plus a set of compiled extensions like xformers, flash-attention and apex. When any two links mismatch, the error usually surfaces somewhere entirely unrelated — which is what makes it so hard to diagnose.

Reproducing someone else's work is worse. A two-year-old paper used the PyTorch and CUDA of its day; running it on a current stack means a chain of API changes and subtly different operator behaviour. What you need there is not the newest environment but the ability to pin an exact one.

Prebuilt images solve the first problem: stable, tested version combinations for the common frameworks, ready to train on boot. Custom images solve the second: point us at any Docker image and reproduce the author's environment exactly, with the whole version matrix under your control.

02 — Workloads

Shapes of training work

From getting one GPU working to scaling out, from reproducing papers to training from scratch.

  • Single-GPU experiments and tuning

    Changing architecture, sweeping hyperparameters, trying loss functions — work that needs fast iteration rather than peak compute. A $0.540/hr 4090 is plenty; destroy it when the sweep ends.

  • Multi-GPU data parallelism

    When the dataset is large and one epoch takes too long, DDP or FSDP splits the batch across cards. Filter the marketplace for multi-GPU nodes, up to 14 per node.

  • Paper reproduction

    Use a custom image to restore the author's environment precisely and avoid result drift from version skew. This is a quiet advantage of renting over a fixed local install.

  • Computer vision and classical models

    Detection, segmentation, super-resolution, point clouds — not every GPU job is an LLM. These have modest VRAM needs, which is exactly where consumer cards win hardest on value.

03 — Stack

All three frameworks, preinstalled

Stock NVIDIA drivers and CUDA, no patched runtime — it behaves the way it does on your workstation.

  • PyTorch

    With matched CUDA and cuDNN, plus torchvision, torchaudio and the usual compiled extensions. Launch distributed jobs with torchrun; DDP and FSDP work out of the box.

    PyTorch · torchrun · DDP/FSDP

  • TensorFlow / Keras

    GPU build installed with device visibility verified, including Keras 3. The right pick for existing TF codebases and SavedModel deployment paths.

    TensorFlow · Keras 3 · SavedModel

  • JAX and distributed extras

    JAX on the CUDA backend with Flax and Optax. DeepSpeed and Accelerate are in the image too, so large-model training needs no extra installation.

    JAX · Flax · DeepSpeed · Accelerate

04 — Choosing a GPU

Sizing a card for training

Training is hungrier than inference — beyond weights you hold gradients, optimiser state and activations. As a rule of thumb it needs three to four times the inference footprint.

Training scenarioVRAM neededCheapest availableNotes
CV models and small networks12GBRTX 3060$0.100/hrDetection, segmentation and classification have modest VRAM needs; consumer cards are entirely sufficient.
Mid-size models, large batches24GBTesla V100$0.188/hrFor jobs that need a larger batch size to converge stably. 24GB is the value sweet spot.
Large-model single-GPU training48GBQ RTX 8000$0.508/hrThe 48GB tier, which with gradient checkpointing trains a substantial model.
Distributed training nodes80GBA100 SXM4$1.088/hrThe 80GB tier at minimum. For multi-GPU, prefer nodes with high-speed interconnect — communication overhead varies enormously.

"Cheapest available" is derived from live inventory — the lowest per-GPU rate among models that clear the VRAM bar — and moves with the market. Multi-GPU nodes rent whole.

05 — Getting started

Your first training run

  • 01

    Pick a framework image

    PyTorch, TensorFlow or JAX. If your code pins a specific version combination, supplying your own Docker image is the safer route.

  • 02

    Get code and data across

    git clone from the web terminal, or scp from your machine. For large datasets, stage them in object storage and pull from inside the instance — the host's bandwidth is far faster.

  • 03

    Train and checkpoint

    Configure periodic checkpointing in your script, download the weights when it finishes, then destroy the instance. Destroying releases the disk; there is no recycle bin.

06 — FAQ

Deep learning training

How much VRAM does training need, and how do I estimate it?

Roughly: training memory equals parameter count times (weights + gradients + optimiser state). With Adam in FP16 that is about 16 bytes per parameter, plus activations. So a 1B-parameter model typically needs 16GB or more. Gradient checkpointing trades compute time for memory and cuts that figure substantially.

How much faster is multi-GPU?

Data-parallel scaling is close to linear in theory but is limited by communication overhead in practice. Within one node, cards talk over PCIe or NVLink and efficiency is good; across nodes it depends on the network. Get it working on one GPU first, then evaluate — on small models communication overhead can eat most of the gain.

Can I pin the CUDA version?

Prebuilt images carry the CUDA build that matches their framework. If you need an exact combination — reproducing a paper, for instance — supply your own Docker image so the whole chain is yours and our image updates cannot affect it.

How do I keep a long training run from being interrupted?

Instances keep running while the balance holds, and low-balance warnings arrive in advance. Configure periodic checkpointing regardless — it is necessary in any environment and especially in the cloud.

What is the fastest way to upload a dataset?

Under a few GB, scp directly. For larger sets, stage them in object storage or any publicly reachable location and pull with wget or aria2 from inside the instance. That uses the host's public bandwidth, typically one to two orders of magnitude faster than uploading from home broadband.

Start building on NexGPU

Enterprise R&D team or solo developer — either way, your first job can be running in minutes.

Sign up to browse live network pricing. No payment method required.