Skip to main content

AI Agents

Run your agents on hardware you control

The expensive thing about an agent is that it thinks repeatedly — dozens of model calls for a single task is normal. Per-token billing makes that unpredictable. Per-GPU-hour billing makes it a fixed number.

Local inference per hour, from
$0.540
Per-token charges
0
Continuous operation
24/7

01 — When it applies

Why agents suit self-hosted inference

An agent's cost structure differs sharply from a chat application's. A conversation is one question and one answer. An agent completing a task may loop a dozen times through planning, tool use, observation and replanning — each loop a full model call, many of them carrying long context.

Under per-token billing that does not grow linearly, it jumps. An agent stuck in a loop can burn a day's budget before anyone notices. Under per-GPU-hour billing, however many times it thinks, the ceiling is known.

The second reason is latency and stability. Every round depends on the previous one, so rate limits, jitter and occasional timeouts from an external API get amplified along the whole chain. On a machine you rented, that class of uncertainty disappears — the only bottleneck left is the card itself.

02 — Workloads

Common agent shapes

What they share: many calls, long context, long runtimes.

  • Research and synthesis

    Give it a topic and let it search, read, cross-check and write a report. These have the highest call counts, which makes them where self-hosted inference saves the most.

  • Code assistants and automation

    Read the repo, locate the problem, edit, run tests, read the results and edit again. The loop count is unpredictable. On local inference you can let it iterate until it is right.

  • Multi-agent collaboration

    CrewAI and AutoGen have several roles converse to advance a task. More roles and more rounds mean a frightening token bill — and an unchanged GPU-hour bill.

  • Long-lived automation

    Monitoring, scheduled analysis, event-driven handling — agents that need to stay online. A resident instance costs a predictable fixed amount per month.

03 — Stack

Agent frameworks and the inference base

Inference runs locally; pick whichever framework you like — most only need an OpenAI-compatible URL.

  • Base: vLLM

    High-frequency agent calls care most about throughput. vLLM's continuous batching lets parallel agents share one card without slowing each other, and serves an OpenAI-compatible endpoint on boot.

    vLLM · OpenAI-compatible · High throughput

  • Orchestration: LangChain / LlamaIndex

    The de facto standard for tool calling, memory and retrieval augmentation. Point base_url at your local vLLM and the rest of the code is untouched.

    LangChain · LlamaIndex · RAG

  • Multi-role: CrewAI / AutoGen

    For complex tasks needing several cooperating agents. More rounds means more calls, which amplifies the cost advantage of local inference.

    CrewAI · AutoGen · Multi-agent

04 — Choosing a GPU

How much card an agent needs

Agents are sensitive to reasoning quality (one bad step derails the chain) and to throughput (many calls). This is not the place to economise hard.

ScaleVRAM neededCheapest availableNotes
Single agent, 7B quantised12GBRTX 3060$0.100/hrEnough to validate an idea or run simple flows. Small models tend to misjudge mid-chain on complex reasoning.
Single agent, 14B-class32GBTesla V100$0.188/hrNoticeably more reliable at tool calling and multi-step reasoning — the practical starting point for self-hosted agents.
Parallel agents, 32B quantised48GBQ RTX 8000$0.508/hrSeveral roles share one card and need more VRAM for KV cache. The balance point between quality and concurrency.
Production, 70B-class80GBA100 SXM4$1.088/hrComplex decision chains need stronger reasoning — one card at the 80GB tier, or a multi-GPU node.

"Cheapest available" is derived from live inventory — the lowest per-GPU rate among models that clear the VRAM bar — and moves with the market. Multi-GPU nodes rent whole.

05 — Getting started

Moving agents onto self-hosted inference

  • 01

    Boot a vLLM instance

    Pick a card that clears the VRAM bar and start the vLLM template. Allocate the port your inference service will expose at order time — mappings cannot be added later.

  • 02

    Change one base_url

    LangChain, CrewAI and AutoGen all accept a custom OpenAI-compatible endpoint. Swap in the instance's public address and the migration is done.

  • 03

    Run first, tune the model later

    Get the flow working on a mid-size model, then decide from observed behaviour whether to move up. Changing models means booting a new instance, not changing code.

06 — FAQ

AI agent deployment

Is self-hosted inference actually cheaper than an API for agents?

It depends on volume. At low frequency a commercial API wins, because you pay nothing while idle. But agents are high-frequency by nature — dozens of calls per task, hundreds of tasks a day. At that scale a $0.540/hr card costs clearly less than the equivalent token bill, and the ceiling is known in advance.

Can a small model carry an agent's reasoning chain?

For simple flows, yes; for complex decisions it struggles. The problem is error accumulation — a wrong judgement at step three derails everything after it. Start at 14B-class and only consider going smaller once the flow is proven.

Agents call external tools — can the instance reach the internet?

Yes, instances have full public network access for external APIs, web fetching or connecting to your own database. Egress is metered per GB at a median around $0.0081/GB, which at normal agent traffic volumes is negligible.

How do I keep an agent running 24/7?

Keep the instance running and the balance funded; low-balance warnings arrive well in advance. Supervise the agent process with systemd or supervisor so a crash restarts automatically.

Can several agents share one inference service?

Yes, and they should. vLLM's continuous batching is built for exactly this — concurrent requests share one card's compute, and total throughput far exceeds giving each agent its own GPU.

Start building on NexGPU

Enterprise R&D team or solo developer — either way, your first job can be running in minutes.

Sign up to browse live network pricing. No payment method required.