AI fine-tuning
Teach a model your vocabulary
A general model does not know your product names, your ticket format or your industry's shorthand. Fine-tuning is the most direct way to press that knowledge into the weights — and the compute it needs is far cheaper rented by the hour than bought.
- 7B QLoRA per hour, from
- $0.723
- Typical QLoRA VRAM bar
- 24GB
- Max VRAM per node
- 2152GB
01 — When it applies
When to fine-tune, and when not to
Start with when not to. If the goal is for the model to know a body of facts — product documentation, policy text, historical records — retrieval augmentation beats fine-tuning: it is cheaper, updates instantly, and is far easier to debug when it is wrong. Fine-tuning solves a different problem: teaching an output format, a tone, or a way of reasoning in a domain.
The signals that you should fine-tune are recognisable. Your prompt has grown long and elaborate and is still inconsistent. You need reliably structured output. The base model keeps mangling your field's terminology. All three are cases where the knowledge needs to live in the weights rather than in the context window.
On cost, LoRA and QLoRA have pushed the barrier very low. A 7B QLoRA run fits on a single 24GB card, takes a few hours, and costs single-digit dollars. That means fine-tuning can be an experiment rather than a project — get the data mix wrong and simply run it again.
02 — Workloads
Common fine-tuning jobs
From adjusting tone to teaching a whole domain's reasoning.
Domain adaptation
Medicine, law, finance, manufacturing — teaching a general model your field's terminology and conventions. A few thousand to a few tens of thousands of good samples usually produces a visible improvement.
Structured output
When you need reliable JSON, SQL or a specific report format. Prompt-level constraints always carry a failure rate; fine-tuning raises format adherence substantially.
Tone and persona
Support assistants, role-play, brand voice. A system prompt gets you most of the way; fine-tuning closes the rest — and it does not consume context length.
Distillation and cost reduction
Train a small model on a large model's outputs to press 70B behaviour into 7B. Inference cost can drop by an order of magnitude, which pays for itself quickly at high concurrency.
03 — Stack
Fine-tuning frameworks, preinstalled
Three directions: least effort, least VRAM, most control.
LLaMA Factory — least effort
An all-in-one tool with a web interface covering hundreds of models, LoRA/QLoRA/full-parameter and SFT/DPO/PPO. Fill in a few fields and it runs — ideal for fast validation.
LLaMA Factory · WebUI · SFT/DPO
Unsloth — least VRAM
Deeply optimised training kernels let the same card handle a larger model or a longer sequence, and run faster doing it. Reach for it first when VRAM is tight.
Unsloth · Low VRAM · Fused kernels
Axolotl — most control
Config-file driven with nearly every training detail exposed and a rich library of community configs. Right when you already know what you want to change.
Axolotl · YAML config · Multi-GPU
04 — Choosing a GPU
Which card for which training job
Fine-tuning VRAM depends on model size, training method and sequence length together. These are practical figures for common combinations.
| Training job | VRAM needed | Cheapest available | Notes |
|---|---|---|---|
| 7B model, QLoRA | 16GB | RTX 4060 Ti$0.126/hr | 4-bit quantised loading — the lowest-VRAM option and the standard entry configuration. |
| 7B model, LoRA (bf16) | 24GB | Tesla V100$0.188/hr | No quantisation, so training is steadier and converges better. Consumer flagships at 24GB cover this exactly. |
| 32B model, QLoRA | 48GB | Q RTX 8000$0.508/hr | Needs the 48GB tier, with extra headroom on top for long sequences. |
| 70B model, QLoRA | 80GB | A100 SXM4$1.088/hr | The 80GB tier at minimum on one card; a multi-GPU node is more comfortable. |
"Cheapest available" is derived from live inventory — the lowest per-GPU rate among models that clear the VRAM bar — and moves with the market. Multi-GPU nodes rent whole.
05 — Getting started
Running a fine-tune
01
Prepare the data before you boot
Fine-tuning succeeds or fails mostly on data. Get the training set into a standard format — usually JSONL instruction/response pairs — and verify it before renting anything. Do not leave a GPU idling while you clean data.
02
Pick a card and an image
Choose the VRAM tier from the table above and boot LLaMA Factory or Axolotl. Size the disk generously: base weights, checkpoints and logs all take space.
03
Train, download, destroy
LoRA weights are typically tens to hundreds of megabytes, so download them directly. Destroy the instance once you have them — a stopped instance keeps billing for disk.
06 — FAQ
LLM fine-tuning
What does one fine-tuning run cost?
LoRA, QLoRA or full-parameter?
How much training data do I need?
What if the instance is destroyed mid-training?
How do I deploy the fine-tuned model?
Related solutions
Other ways to use it
Same compute network — swap the image and it becomes a different production line.
AI Agents
Deploy and scale agents on LangChain, CrewAI
Private LLM Deployment
Open weights, running on your own machine
AI Image & Video
Stable Diffusion, FLUX and ComfyUI, ready to run
AI Text Generation
vLLM, TGI and Ollama, live in minutes
AI/ML Frameworks
Native PyTorch, TensorFlow and JAX
Audio to Text
GPU-accelerated Whisper transcription
Start building on NexGPU
Enterprise R&D team or solo developer — either way, your first job can be running in minutes.
Sign up to browse live network pricing. No payment method required.
