Skip to main content

AI fine-tuning

Teach a model your vocabulary

A general model does not know your product names, your ticket format or your industry's shorthand. Fine-tuning is the most direct way to press that knowledge into the weights — and the compute it needs is far cheaper rented by the hour than bought.

7B QLoRA per hour, from
$0.723
Typical QLoRA VRAM bar
24GB
Max VRAM per node
2152GB

01 — When it applies

When to fine-tune, and when not to

Start with when not to. If the goal is for the model to know a body of facts — product documentation, policy text, historical records — retrieval augmentation beats fine-tuning: it is cheaper, updates instantly, and is far easier to debug when it is wrong. Fine-tuning solves a different problem: teaching an output format, a tone, or a way of reasoning in a domain.

The signals that you should fine-tune are recognisable. Your prompt has grown long and elaborate and is still inconsistent. You need reliably structured output. The base model keeps mangling your field's terminology. All three are cases where the knowledge needs to live in the weights rather than in the context window.

On cost, LoRA and QLoRA have pushed the barrier very low. A 7B QLoRA run fits on a single 24GB card, takes a few hours, and costs single-digit dollars. That means fine-tuning can be an experiment rather than a project — get the data mix wrong and simply run it again.

02 — Workloads

Common fine-tuning jobs

From adjusting tone to teaching a whole domain's reasoning.

  • Domain adaptation

    Medicine, law, finance, manufacturing — teaching a general model your field's terminology and conventions. A few thousand to a few tens of thousands of good samples usually produces a visible improvement.

  • Structured output

    When you need reliable JSON, SQL or a specific report format. Prompt-level constraints always carry a failure rate; fine-tuning raises format adherence substantially.

  • Tone and persona

    Support assistants, role-play, brand voice. A system prompt gets you most of the way; fine-tuning closes the rest — and it does not consume context length.

  • Distillation and cost reduction

    Train a small model on a large model's outputs to press 70B behaviour into 7B. Inference cost can drop by an order of magnitude, which pays for itself quickly at high concurrency.

03 — Stack

Fine-tuning frameworks, preinstalled

Three directions: least effort, least VRAM, most control.

  • LLaMA Factory — least effort

    An all-in-one tool with a web interface covering hundreds of models, LoRA/QLoRA/full-parameter and SFT/DPO/PPO. Fill in a few fields and it runs — ideal for fast validation.

    LLaMA Factory · WebUI · SFT/DPO

  • Unsloth — least VRAM

    Deeply optimised training kernels let the same card handle a larger model or a longer sequence, and run faster doing it. Reach for it first when VRAM is tight.

    Unsloth · Low VRAM · Fused kernels

  • Axolotl — most control

    Config-file driven with nearly every training detail exposed and a rich library of community configs. Right when you already know what you want to change.

    Axolotl · YAML config · Multi-GPU

04 — Choosing a GPU

Which card for which training job

Fine-tuning VRAM depends on model size, training method and sequence length together. These are practical figures for common combinations.

Training jobVRAM neededCheapest availableNotes
7B model, QLoRA16GBRTX 4060 Ti$0.126/hr4-bit quantised loading — the lowest-VRAM option and the standard entry configuration.
7B model, LoRA (bf16)24GBTesla V100$0.188/hrNo quantisation, so training is steadier and converges better. Consumer flagships at 24GB cover this exactly.
32B model, QLoRA48GBQ RTX 8000$0.508/hrNeeds the 48GB tier, with extra headroom on top for long sequences.
70B model, QLoRA80GBA100 SXM4$1.088/hrThe 80GB tier at minimum on one card; a multi-GPU node is more comfortable.

"Cheapest available" is derived from live inventory — the lowest per-GPU rate among models that clear the VRAM bar — and moves with the market. Multi-GPU nodes rent whole.

05 — Getting started

Running a fine-tune

  • 01

    Prepare the data before you boot

    Fine-tuning succeeds or fails mostly on data. Get the training set into a standard format — usually JSONL instruction/response pairs — and verify it before renting anything. Do not leave a GPU idling while you clean data.

  • 02

    Pick a card and an image

    Choose the VRAM tier from the table above and boot LLaMA Factory or Axolotl. Size the disk generously: base weights, checkpoints and logs all take space.

  • 03

    Train, download, destroy

    LoRA weights are typically tens to hundreds of megabytes, so download them directly. Destroy the instance once you have them — a stopped instance keeps billing for disk.

06 — FAQ

LLM fine-tuning

What does one fine-tuning run cost?

A 7B QLoRA run over a few thousand samples for three epochs typically takes two to six hours. At $0.723 per hour, that is under ten dollars for a complete run — cheap enough to treat fine-tuning as an experiment and simply rerun it with a different data mix.

LoRA, QLoRA or full-parameter?

LoRA or QLoRA in most cases. They train a small set of added parameters, so VRAM is low, training is fast, and the output weights are tens of megabytes — with results close to full fine-tuning on most tasks. Full-parameter training needs several times the VRAM and time, and is only worth it for deep domain adaptation or distillation.

How much training data do I need?

It depends on the task. Changing output format or tone often works with several hundred to a thousand good samples; domain adaptation usually wants thousands to tens of thousands. Quality matters far more than volume — a thousand carefully labelled samples beats ten thousand noisy ones.

What if the instance is destroyed mid-training?

It will not be while your balance holds. Low-balance warning emails go out well ahead, and only an exhausted balance triggers an automatic stop. Fund the account before a long run and configure periodic checkpointing in your training script — which you should be doing locally too.

How do I deploy the fine-tuned model?

vLLM can load LoRA weights directly, or you can merge them into the base model first. See our LLM inference deployment page — same account, same marketplace, different image.

Start building on NexGPU

Enterprise R&D team or solo developer — either way, your first job can be running in minutes.

Sign up to browse live network pricing. No payment method required.