Refurbished GPU Servers for AI Inference: A Right-Sizing Guide for Indian Teams
Quick answer: an AI inference server needs enough GPU memory (VRAM) to hold the model's weights plus its working memory, and enough throughput to serve your users — it does not need the memory and multi-GPU interconnect that training demands. As a rule, a model that fits on a single GPU is cheaper and simpler to serve than one split across several. That is why a 24GB RTX 4090, a 48GB RTX A6000 or L40S, or an 80GB A100 — new or refurbished — is often enough for inference, where training the same model would need a much larger system. Match the GPU to the model size and the number of concurrent users, not to the biggest number on a spec sheet.
Why inference needs a different server from training
Training a model keeps far more in GPU memory than serving it. Alongside the weights, training holds gradients, optimizer states and activations; a commonly cited figure for mixed-precision training with the Adam optimizer is around 16 bytes of memory per parameter. Inference holds the weights, a working cache and some runtime overhead, and no optimizer at all. Inference also runs continuously and answers many small requests, so what matters is latency and steady throughput rather than the raw multi-GPU scale that training rewards.
In practice that means an inference server can usually be smaller, cheaper and easier to run than a training node — provided you size the GPU memory correctly. Undersize the VRAM and the model will not load; oversize it and you pay for memory you never use.
Step 1: Work out how much VRAM your model needs
The starting point is simple arithmetic: memory for weights = number of parameters × bytes per parameter. Half precision (FP16 or BF16) uses 2 bytes per parameter, 8-bit quantization uses about 1 byte, and 4-bit quantization uses about 0.5 byte. The table shows weights only — the real requirement is higher.
| Model size | FP16 / BF16 weights | 8-bit weights | 4-bit weights |
|---|---|---|---|
| 7B parameters | ~14 GB | ~7 GB | ~3.5 GB |
| 13B parameters | ~26 GB | ~13 GB | ~6.5 GB |
| 34B parameters | ~68 GB | ~34 GB | ~17 GB |
| 70B parameters | ~140 GB | ~70 GB | ~35 GB |
On top of the weights you need room for the KV cache, the memory a language model uses to remember the conversation so far. It grows with the length of the context and with the number of requests served at the same time, so a model that just fits on paper can run out of memory under real traffic. Leave headroom, and test with your own prompt lengths and concurrency before you commit to hardware. Quantization (8-bit or 4-bit) is the usual way to fit a bigger model on a smaller GPU, at some cost in output quality that you should measure for your use case.
Step 2: Match the model to a GPU
Serverwale's GPU servers range covers the cards below. VRAM sizes and starting prices are from our GPU servers page; prices are indicative and depend on configuration and stock.
| GPU | VRAM | Our positioning | Starting price | Fits on one GPU (weights only) |
|---|---|---|---|---|
| NVIDIA RTX 4090 | 24 GB | Workstation AI and rendering; value for inference and fine-tuning | From ₹2,80,000 | 7B in FP16; 13B in 8-bit; ~34B in 4-bit |
| NVIDIA V100 (refurbished) | 32 GB | Lower-cost option, best for CNN and computer-vision workloads | From ₹4,50,000 | 13B in 8-bit; ~34B in 4-bit |
| NVIDIA RTX A6000 | 48 GB | AI plus 3D rendering | From ₹4,80,000 | 13B in FP16; 34B in 8-bit; 70B in 4-bit |
| NVIDIA L40S | 48 GB | AI inference and VDI | On request | 13B in FP16; 34B in 8-bit; 70B in 4-bit |
| NVIDIA A100 | 80 GB | AI training and inference | From ₹18,00,000 | 34B in FP16; 70B in 8-bit (little headroom) |
| NVIDIA H100 | 80 GB | LLM training and generative AI | On request | Same memory as A100, with higher throughput |
A 70B model in FP16 (~140 GB of weights) does not fit on a single 80GB card and needs at least two GPUs. "Fits" here means the weights load; the KV cache still has to fit in what is left, so the tightest fits in the table need short contexts or few concurrent users.
Step 3: Size for throughput, not just for fit
Once the model loads, the question is how many users it can serve. Three things decide that:
- Concurrency. More simultaneous requests means a larger KV cache and more compute. Serving software such as vLLM, NVIDIA Triton or Ollama batches requests to use the GPU efficiently; the right choice depends on your model and team.
- Latency target. A chatbot answering in real time needs a different setup from an overnight batch job summarizing documents. Batch workloads can run on a smaller GPU because nobody is waiting.
- Multiple small models. If you run several small services rather than one large model, a data-centre card like the A100 can be partitioned into isolated slices (Multi-Instance GPU), which helps you use one card for several workloads.
The CPU, system RAM and storage around the GPU matter less for inference than for training, but they should not starve it: fast NVMe storage loads model weights quickly, and a balanced host platform avoids bottlenecks — our Xeon vs EPYC comparison helps you pick the CPU that feeds the GPUs.
Where refurbished GPU servers make sense for inference
Inference is a steady, always-on workload, which favours owning hardware once your usage is predictable. Refurbished data-centre GPUs cost less than new ones — Serverwale lists refurbished V100 and A100 systems at 50–60% below new — and each system is 72-point tested with CUDA, PyTorch or TensorFlow set up and up to a 3-year warranty.
Be clear about the trade-offs. Older architectures can lack support for newer number formats and features: for example, V100 predates BF16 support, so it runs models in FP16, and the newest low-precision formats need newer GPUs. Check that your inference framework and model format support the card before you buy. For the newest generation, or if your software depends on those features, a new GPU or a custom ProStation build may be the better fit. Also compare running costs: GPUs draw significant power and need cooling and a UPS, so plan the room as well as the server (see our UPS sizing guide).
Buy, rent or use the cloud?
If your inference load is steady all day, buying usually costs less over time. If you are running a pilot or a short project, renting avoids the capital outlay — see our high-end GPU server rental page and the general rent vs own analysis. Keeping inference on your own hardware also keeps prompts and data inside your premises, which matters for regulated or sensitive workloads; our AI infrastructure page and the on-premise AI infrastructure guide cover the wider build. For a look at the cloud alternative, ProStation Systems compares on-premise AI servers and cloud GPUs. For approximate GPU prices, see the GPU server price guide, and for choosing between the popular cards, our A100 vs RTX A6000 vs RTX 4090 comparison.
Checklist before you order an inference GPU server
- Note the model, its parameter count and the precision or quantization you plan to use.
- Estimate weights from the table, then add headroom for the KV cache at your real context length and concurrency.
- Set your latency target and expected number of simultaneous users.
- Confirm your serving software supports the GPU generation you are considering.
- Plan power, cooling and UPS for a machine that runs around the clock.
- After delivery, run the checks in our post-delivery testing checklist and put the server under a server AMC if it will run production traffic.
Frequently Asked Questions
How much GPU memory do I need to run a 7B or 13B model for inference?
Weights alone need about 14 GB for a 7B model and 26 GB for a 13B model in FP16, or roughly half that in 8-bit. Add memory for the KV cache and runtime overhead, which grows with context length and concurrent users. A 24GB RTX 4090 suits a 7B model in FP16; a 13B model in FP16 is better placed on a 48GB card.
Is an inference server cheaper than a training server?
Usually yes. Inference does not store gradients or optimizer states, so the same model needs much less GPU memory, and it can often run on a single GPU instead of a multi-GPU cluster. The exact saving depends on the model size and how many users you serve.
Which GPU is best for LLM inference in India — RTX 4090, A6000, L40S or A100?
It depends on model size and load. The 24GB RTX 4090 is the value option for smaller models; the 48GB RTX A6000 and L40S handle larger models and quantized 70B models; the 80GB A100 is for the largest single-GPU models and heavier concurrency. Choose by the VRAM your model needs first, then by throughput.
Are refurbished GPU servers reliable enough for production inference?
Refurbished data-centre GPUs are designed for continuous duty. Serverwale tests each GPU system with a 72-point check and provides up to a 3-year warranty; for production traffic, pair it with a server AMC and a standby plan for critical services. Confirm software support for the GPU generation before buying.
Should I rent or buy a GPU server for inference?
Buy when the workload is steady and runs most of the day; rent for pilots, short projects or spikes. Serverwale offers GPU servers on monthly rental from a one-month minimum, so you can test a configuration before committing to a purchase.
Can one GPU serve several models at the same time?
Yes, if the models and their caches fit in the GPU's memory together. Data-centre cards such as the A100 can also be partitioned into isolated instances (Multi-Instance GPU), which suits several small services on one card.
Talk to Serverwale about your inference server
Tell us the model, the precision and the number of users, and we will recommend a GPU configuration, new or refurbished, with pan-India delivery. Call or WhatsApp +91-87962-44410, or contact us for a quote on your GPU server.

