Introduction
Linux
Understanding vLLM
vLLM is an open-source software tool designed to run Large Language Models (LLMs) at high speed with maximum efficiency. When you run AI models on normal setup tools, generating text can be slow and take up huge amounts of graphics card memory. vLLM solves this issue by optimizing how the computer processes text, allowing you to serve AI models much faster while using far less memory.
At the heart of vLLM is a technology called PagedAttention. Traditional AI tools hold onto large, fixed blocks of memory for every request, which wastes valuable GPU resources. PagedAttention works similarly to how modern operating systems manage computer memory by breaking data into tiny pages. This allows vLLM to use almost 100% of your GPU memory efficiently, letting your computer handle many user requests at the same time without slowing down or running out of memory.
Another major advantage of vLLM is its built-in server that mimics the official OpenAI API structure. This means if you already have applications built to work with OpenAI, you can point them to your self-hosted vLLM server without rewriting your code. You get full control over your data, lower operating costs, and complete privacy while maintaining industry-standard compatibility.
Prerequisites
- Linux Operating System: vLLM runs best on Linux distributions such as Ubuntu 22.04 or higher.
- NVIDIA GPU: A dedicated NVIDIA graphics card with CUDA support (CUDA 12.0 or newer) and updated NVIDIA drivers.
- Python: Python version 3.10 to 3.13 installed on your machine (Python 3.12 is recommended).
- Internet Connection: A stable internet connection to download model files and Python packages.
- Basic Terminal Knowledge: Basic comfort with running commands in a Linux terminal.
Step-by-Step Installation
Install uv
uv is an extremely fast Python package manager recommended by the vLLM team to handle installations smoothly.
uv:
# Download and install uv curl -LsSf https://astral.sh/uv/install.sh | sh # Update your system PATH to use uv immediately source ~/.bashrc
Create a Virtual Environment
# Create a virtual environment with Python 3.12 uv venv --python 3.12 --seed # Activate the virtual environment source .venv/bin/activate
Install vLLM
uv. The --torch-backend=auto flag automatically checks your system's GPU driver and installs the matching PyTorch version.
# Install vLLM automatically matched to your GPU driver uv pip install vllm --torch-backend=auto
use uv pip install vllm --extra-index-url [https://wheels.vllm.ai/rocm/] (https://wheels.vllm.ai/rocm/))
Start the vLLM API Server
# Launch the vLLM server vllm serve Qwen/Qwen2.5-1.5B-Instruct
http://localhost:8000.
Test Your Server
curl http://localhost:8000/v1/models
curl http://localhost:8000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "Qwen/Qwen2.5-1.5B-Instruct",
"messages": [
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": "Explain what a vLLM server is in one simple sentence."}
],
"max_tokens": 50,
"temperature": 0.7
}'
Querying Your Server with Python (Optional)
openai library:
uv pip install openai
test_server.py and add the following code:
from openai import OpenAI
# Connect to your local vLLM server
client = OpenAI(
api_key="EMPTY", # vLLM does not require a key by default
base_url="http://localhost:8000/v1",
)
# Request a response from your local model
chat_response = client.chat.completions.create(
model="Qwen/Qwen2.5-1.5B-Instruct",
messages=[
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": "What are the main benefits of using vLLM?"},
],
)
print(chat_response.choices[0].message.content)
python test_server.py
CTCservers Recommended Tutorials
Web, Network
Step-by-Step Guide: Install AMD ROCm on Ubuntu with RX 6600 GPU
Learn how to quickly and easily set up AMD ROCm on Ubuntu for your RX 6600 GPU, enabling powerful machine learning, AI workloads, and GPU-accelerated computing right on your system.
Web, Network, Linux, Mysql, Ubuntu
LAMP Setup Guide 2026: Ubuntu & Debian | CTCservers
Install a secure LAMP stack on Debian or Ubuntu. Follow our step-by-step guide to configure Linux, Apache, MySQL, and PHP for your web server.
Web, Network, Ubuntu
Deploy Phi-3 with Ollama on Ubuntu GPU | CTCservers
Learn how to easily deploy the Phi-3 LLM on an Ubuntu 24.04 GPU server using Ollama and WebUI. Follow our step-by-step tutorial for seamless AI hosting.
Discover CTCservers Dedicated Server Locations
CTCservers servers are available around the world, providing diverse options for hosting websites. Each region offers unique advantages, making it easier to choose a location that best suits your specific hosting needs.