The Ultimate Beginner’s Guide to Setting Up a vLLM Server

This guide will walk you through setting up your own high-speed AI inference server using vLLM in just a few simple steps.

seafile

Introduction

Linux

Understanding vLLM

vLLM is an open-source software tool designed to run Large Language Models (LLMs) at high speed with maximum efficiency. When you run AI models on normal setup tools, generating text can be slow and take up huge amounts of graphics card memory. vLLM solves this issue by optimizing how the computer processes text, allowing you to serve AI models much faster while using far less memory.

At the heart of vLLM is a technology called PagedAttention. Traditional AI tools hold onto large, fixed blocks of memory for every request, which wastes valuable GPU resources. PagedAttention works similarly to how modern operating systems manage computer memory by breaking data into tiny pages. This allows vLLM to use almost 100% of your GPU memory efficiently, letting your computer handle many user requests at the same time without slowing down or running out of memory.

Another major advantage of vLLM is its built-in server that mimics the official OpenAI API structure. This means if you already have applications built to work with OpenAI, you can point them to your self-hosted vLLM server without rewriting your code. You get full control over your data, lower operating costs, and complete privacy while maintaining industry-standard compatibility.

Prerequisites

  • Linux Operating System: vLLM runs best on Linux distributions such as Ubuntu 22.04 or higher.
  • NVIDIA GPU: A dedicated NVIDIA graphics card with CUDA support (CUDA 12.0 or newer) and updated NVIDIA drivers.
  • Python: Python version 3.10 to 3.13 installed on your machine (Python 3.12 is recommended).
  • Internet Connection: A stable internet connection to download model files and Python packages.
  • Basic Terminal Knowledge: Basic comfort with running commands in a Linux terminal.

Step-by-Step Installation

1

Install uv

uv is an extremely fast Python package manager recommended by the vLLM team to handle installations smoothly.
Run the following command to install uv:
BASH
# Download and install uv
curl -LsSf https://astral.sh/uv/install.sh | sh

# Update your system PATH to use uv immediately
source ~/.bashrc
2

Create a Virtual Environment

A virtual environment keeps your project dependencies clean and separate from your system files.
Run these commands to create and enter a Python 3.12 environment:
BASH
# Create a virtual environment with Python 3.12
uv venv --python 3.12 --seed

# Activate the virtual environment
source .venv/bin/activate
3

Install vLLM

Install vLLM using uv. The --torch-backend=auto flag automatically checks your system's GPU driver and installs the matching PyTorch version.
Bash
# Install vLLM automatically matched to your GPU driver
uv pip install vllm --torch-backend=auto
(Note: If you are using an AMD ROCm GPU instead of NVIDIA, use uv pip install vllm --extra-index-url [https://wheels.vllm.ai/rocm/] (https://wheels.vllm.ai/rocm/))
4

Start the vLLM API Server

Now you can start your local server. We will use Qwen2.5-1.5B-Instruct, a lightweight and fast AI model. vLLM will automatically download the model files from Hugging Face the first time you run it.
Bash
# Launch the vLLM server
vllm serve Qwen/Qwen2.5-1.5B-Instruct
Keep this terminal window open. Your server is now running at http://localhost:8000.
5

Test Your Server

Open a new terminal window to test if your server is responding.
1. Check if the model is active:
Bash
curl http://localhost:8000/v1/models
2. Send a prompt to your model:
Bash
curl http://localhost:8000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "Qwen/Qwen2.5-1.5B-Instruct",
    "messages": [
      {"role": "system", "content": "You are a helpful assistant."},
      {"role": "user", "content": "Explain what a vLLM server is in one simple sentence."}
    ],
    "max_tokens": 50,
    "temperature": 0.7
  }'
6

Querying Your Server with Python (Optional)

If you want to talk to your server using Python code, install the official openai library:
Bash
uv pip install openai
Create a file named test_server.py and add the following code:
Bash
from openai import OpenAI

# Connect to your local vLLM server
client = OpenAI(
    api_key="EMPTY",  # vLLM does not require a key by default
    base_url="http://localhost:8000/v1",
)

# Request a response from your local model
chat_response = client.chat.completions.create(
    model="Qwen/Qwen2.5-1.5B-Instruct",
    messages=[
        {"role": "system", "content": "You are a helpful assistant."},
        {"role": "user", "content": "What are the main benefits of using vLLM?"},
    ],
)

print(chat_response.choices[0].message.content)
Run the script in your terminal:
Bash
python test_server.py

Discover CTCservers Dedicated Server Locations

CTCservers servers are available around the world, providing diverse options for hosting websites. Each region offers unique advantages, making it easier to choose a location that best suits your specific hosting needs.