AI Agents
Run your agents on hardware you control
The expensive thing about an agent is that it thinks repeatedly — dozens of model calls for a single task is normal. Per-token billing makes that unpredictable. Per-GPU-hour billing makes it a fixed number.
- Local inference per hour, from
- $0.540
- Per-token charges
- 0
- Continuous operation
- 24/7
01 — When it applies
Why agents suit self-hosted inference
An agent's cost structure differs sharply from a chat application's. A conversation is one question and one answer. An agent completing a task may loop a dozen times through planning, tool use, observation and replanning — each loop a full model call, many of them carrying long context.
Under per-token billing that does not grow linearly, it jumps. An agent stuck in a loop can burn a day's budget before anyone notices. Under per-GPU-hour billing, however many times it thinks, the ceiling is known.
The second reason is latency and stability. Every round depends on the previous one, so rate limits, jitter and occasional timeouts from an external API get amplified along the whole chain. On a machine you rented, that class of uncertainty disappears — the only bottleneck left is the card itself.
02 — Workloads
Common agent shapes
What they share: many calls, long context, long runtimes.
Research and synthesis
Give it a topic and let it search, read, cross-check and write a report. These have the highest call counts, which makes them where self-hosted inference saves the most.
Code assistants and automation
Read the repo, locate the problem, edit, run tests, read the results and edit again. The loop count is unpredictable. On local inference you can let it iterate until it is right.
Multi-agent collaboration
CrewAI and AutoGen have several roles converse to advance a task. More roles and more rounds mean a frightening token bill — and an unchanged GPU-hour bill.
Long-lived automation
Monitoring, scheduled analysis, event-driven handling — agents that need to stay online. A resident instance costs a predictable fixed amount per month.
03 — Stack
Agent frameworks and the inference base
Inference runs locally; pick whichever framework you like — most only need an OpenAI-compatible URL.
Base: vLLM
High-frequency agent calls care most about throughput. vLLM's continuous batching lets parallel agents share one card without slowing each other, and serves an OpenAI-compatible endpoint on boot.
vLLM · OpenAI-compatible · High throughput
Orchestration: LangChain / LlamaIndex
The de facto standard for tool calling, memory and retrieval augmentation. Point base_url at your local vLLM and the rest of the code is untouched.
LangChain · LlamaIndex · RAG
Multi-role: CrewAI / AutoGen
For complex tasks needing several cooperating agents. More rounds means more calls, which amplifies the cost advantage of local inference.
CrewAI · AutoGen · Multi-agent
04 — Choosing a GPU
How much card an agent needs
Agents are sensitive to reasoning quality (one bad step derails the chain) and to throughput (many calls). This is not the place to economise hard.
| Scale | VRAM needed | Cheapest available | Notes |
|---|---|---|---|
| Single agent, 7B quantised | 12GB | RTX 3060$0.100/hr | Enough to validate an idea or run simple flows. Small models tend to misjudge mid-chain on complex reasoning. |
| Single agent, 14B-class | 32GB | Tesla V100$0.188/hr | Noticeably more reliable at tool calling and multi-step reasoning — the practical starting point for self-hosted agents. |
| Parallel agents, 32B quantised | 48GB | Q RTX 8000$0.508/hr | Several roles share one card and need more VRAM for KV cache. The balance point between quality and concurrency. |
| Production, 70B-class | 80GB | A100 SXM4$1.088/hr | Complex decision chains need stronger reasoning — one card at the 80GB tier, or a multi-GPU node. |
"Cheapest available" is derived from live inventory — the lowest per-GPU rate among models that clear the VRAM bar — and moves with the market. Multi-GPU nodes rent whole.
05 — Getting started
Moving agents onto self-hosted inference
01
Boot a vLLM instance
Pick a card that clears the VRAM bar and start the vLLM template. Allocate the port your inference service will expose at order time — mappings cannot be added later.
02
Change one base_url
LangChain, CrewAI and AutoGen all accept a custom OpenAI-compatible endpoint. Swap in the instance's public address and the migration is done.
03
Run first, tune the model later
Get the flow working on a mid-size model, then decide from observed behaviour whether to move up. Changing models means booting a new instance, not changing code.
06 — FAQ
AI agent deployment
Is self-hosted inference actually cheaper than an API for agents?
Can a small model carry an agent's reasoning chain?
Agents call external tools — can the instance reach the internet?
How do I keep an agent running 24/7?
Can several agents share one inference service?
Related solutions
Other ways to use it
Same compute network — swap the image and it becomes a different production line.
Private LLM Deployment
Open weights, running on your own machine
AI Fine-tuning
LoRA, QLoRA and full fine-tuning, on demand
AI Image & Video
Stable Diffusion, FLUX and ComfyUI, ready to run
AI Text Generation
vLLM, TGI and Ollama, live in minutes
AI/ML Frameworks
Native PyTorch, TensorFlow and JAX
Audio to Text
GPU-accelerated Whisper transcription
Start building on NexGPU
Enterprise R&D team or solo developer — either way, your first job can be running in minutes.
Sign up to browse live network pricing. No payment method required.
