Installing Prerequisites
We are installing on a very fresh install of Ubuntu 22.04 with an Nvidia 3070ti GPU. Nvidia Drivers were installed at build time. However lets check to see what version of the Nvidia driver we are currently running.
$ nvidia-detector
nvidia-driver-570
While we are at it, let’s make sure that we have nvidia-smi installed which is included in the package shown below
$ sudo apt install nvidia-utils-570
We will need gcc, which the meta-package below will include (along with make and other dev tools).
# sudo apt install build-essential -y
Lets check what version of python we have installed.
$ python3 --version
Python 3.10.12
Now install python-venv which is the python module that supports creating virtual environments. Reference the python version shown in the output in the step above
# sudo apt install python3.10-venv -y
Installing the Nvidia Cuda Toolkit
Note: If you installed the nvidia-cuda-toolkit from default ubuntu noble repo, uninstall it first as that version is probably too old. Uninstall with “sudo apt purge nvidia-cuda* -y“
“The NVIDIA CUDA Toolkit is a software development kit that helps users create GPU-accelerated applications. It includes libraries, compilers, debuggers, and optimization tools”. Since we have an Nvidia GPU we will install it as shown below.
# wget https://developer.download.nvidia.com/compute/cuda/repos/ubuntu2204/x86_64/cuda-ubuntu2204.pin
First we wget the Nvidia Cuda Repository pin as shown above. Then we move it to /etc/apt/preferences.d/ (shown below).
This is a configuration file that is used to prioritize packages from the NVIDIA CUDA repository when installing CUDA on a Linux system. Basically it tells apt where to get the Cuda Toolkit.
$ sudo mv cuda-ubuntu2204.pin /etc/apt/preferences.d/cuda-repository-pin-600
Now we wget the repo. Note that this is a pretty hefty .deb and may take a few minutes.
$ wget https://developer.download.nvidia.com/compute/cuda/12.8.0/local_installers/cuda-repo-ubuntu2204-12-8-local_12.8.0-570.86.10-1_amd64.deb
And install it.
$ sudo dpkg -i cuda-repo-ubuntu2204-12-8-local_12.8.0-570.86.10-1_amd64.deb
$ sudo cp /var/cuda-repo-ubuntu2204-12-8-local/cuda-*-keyring.gpg /usr/share/keyrings/
Finally we install the toolkit.
$ sudo apt-get update
$ sudo apt-get -y install cuda-toolkit-12-8
$ sudo apt install cuda-runtime-12-8
Nvidia states that you need to add the following env vars to your .bashrc, however you will also need to add them to your python virtual env — in venv/bin/activate. These variables will be needed during the install process below.
export CUDA_HOME=/usr/local/cuda
export PATH=${CUDA_HOME}/bin:${PATH}
export LD_LIBRARY_PATH=${CUDA_HOME}/lib64:$LD_LIBRARY_PATH
Installing Instructlab
First we make a directory were we want our python virtual env to live.
$ mkdir instructlab
$ cd instructlab/
Now we create the python virtual environment.
$ python3 -m venv --upgrade-deps venv
$ source venv/bin/activate
Note that every time you want to use instructlab cli (ilab) post-install you must source the file above. venv/bin/activate
The install guide found here has you run this command to clear out the pip package cache (for this one package. Not required for first install, but may be on subsequent install runs.
$ pip cache remove llama_cpp_python
We then run the 3 commands shown below. Apparently there is a known bug and we will not pick up these required packages as needed, so we need to install them manually, and in the order shown below.
$ pip install --upgrade pip wheel setuptools
$ pip install torch
$ pip install flash-attn
Now we are ready to install instructlab.
$ pip install 'instructlab[cuda]' /
-C cmake.args="-DLLAMA_CUDA=on" /
-C cmake.args="-DLLAMA_NATIVE=off"
Lets verify the install
$ ilab --version
ilab, version 0.24.0
Installing vLLM
The instructions have you wrap up the of Instructlab with a pip install of vllm, however this install will fail due as it cannot find nvcc without setting the env variables. So set the following env vars as shown below.
$ export CUDA_HOME=/usr/local/cuda
$ export PATH=${CUDA_HOME}/bin:${PATH}
$ export LD_LIBRARY_PATH=${CUDA_HOME}/lib64:$LD_LIBRARY_PATH
The instructions for installing instructlab with Nvidia Cuda support as documented on instructlab.ai currently have you pip install a very specific version of vllm (with cuda support), however I have found that this command will fail (at least on Fedora 40, Ubuntu 22.04, and Ubuntu 24.04) that command is shown below.
$ pip install vllm@git+https://github.com/opendatahub-io/vllm@2024.08.01
So this is a method that I have found to work around the issue.
First run pip install on vllm
$ pip install vllm
Confirm vllm installed correctly.
$ pip show vllm
Name: vllm
Version: 0.7.3
Summary: A high-throughput and memory-efficient inference and serving engine for LLMs
Home-page: https://github.com/vllm-project/vllm
Author: vLLM Team
Author-email:
License: Apache License
Author-email:
License: Apache 2.0
Location: /home/cpaquin/Workspace/instructlab/venv/lib/python3.10/site-packages
Requires: aiohttp, cmake, fastapi, filelock, lm-format-enforcer, ninja, numpy, nvidia-ml-py, openai, outlines, pillow, prometheus-fastapi-instrumentator, prometheus_client, psutil, py-cpuinfo, pydantic, pyzmq, requests, sentencepiece, tiktoken, tokenizers, torch, torchvision, tqdm, transformers, typing_extensions, uvicorn, vllm-flash-attn, xformers
Now remove vllm
$ pip uninstall vllm
Now run the original pip install command as documented
$ pip install vllm@git+https://github.com/opendatahub-io/vllm@2024.08.01
Now check vllm version, which we can see has cuda support
$ pip show vllm
Name: vllm
Version: 0.5.2.3+cu128
Summary: A high-throughput and memory-efficient inference and serving engine for LLMs
Home-page: https://github.com/vllm-project/vllm
Author: vLLM Team
Author-email:
License: Apache 2.0
Location: /home/cpaquin/Workspace/instructlab/venv/lib/python3.10/site-packages
Requires: aiohttp, cmake, fastapi, filelock, lm-format-enforcer, ninja, numpy, nvidia-ml-py, openai, outlines, pillow, prometheus-fastapi-instrumentator, prometheus_client, psutil, py-cpuinfo, pydantic, pyzmq, requests, sentencepiece, tiktoken, tokenizers, torch, torchvision, tqdm, transformers, typing_extensions, uvicorn, vllm-flash-attn, xformers
Required-by:
Testing VLLM
Below is the command that I ran to test that vllm is working and is in fact, leveraging the GPU
$ vllm serve "TinyLlama/TinyLlama-1.1B-Chat-v1.0" --dtype float16 --gpu-memory-utilization 0.8
One up and running you will see output similar to what is shown below.
INFO: Started server process [1252947]
INFO: Waiting for application startup.
INFO: Application startup complete.
INFO: Uvicorn running on http://0.0.0.0:8000 (Press CTRL+C to quit)
Testing Pytorch Module
You can create and run the python script below to test if your GPU is identified correctly. This step is not required, but its handy in identifying potential issues with gpu/cuda.
import torch
print("CUDA available:", torch.cuda.is_available())
print("CUDA version:", torch.version.cuda)
print("Torch version:", torch.__version__)
print("GPU count:", torch.cuda.device_count())
print("GPU name:", torch.cuda.get_device_name(0) if torch.cuda.is_available() else "None")
Output below for reference.
$ python3 ./test.py
CUDA available: True
CUDA version: 12.4
Torch version: 2.5.1+cu124
GPU count: 1
GPU name: NVIDIA GeForce RTX 3070 Ti
Instructlab Init and Setup
Now that we have ilab installed we can initialize our environment with the command below. This command will download the Taxonomy repository to our local machine along with a config file. Apparently ilab init has a few default GPUs available to choose from, but since my current GPU is

Our config file is located in our home directory, if you want to take a look. Note that our GPU, being a Nvidia RTX 3070 TI, will have differing config values than what we selected above (A100), so in a later step we will need to edit our config
/home/cpaquin/.config/instructlab/config.yaml
To see the info regarding your workstation, you can run the command below. We will use some of the data below to update our config
ilab system info
Platform:
sys.version: 3.10.12 (main, Jan 17 2025, 14:35:34) [GCC 11.4.0]
sys.platform: linux
os.name: posix
platform.release: 6.8.0-52-generic
platform.machine: x86_64
platform.node: scar-Z97-HD3
platform.python_version: 3.10.12
os-release.ID: ubuntu
os-release.VERSION_ID: 22.04
os-release.PRETTY_NAME: Ubuntu 22.04.5 LTS
memory.total: 31.30 GB
memory.available: 27.11 GB
memory.used: 3.60 GB
InstructLab:
instructlab.version: 0.24.0
instructlab-dolomite.version: 0.2.0
instructlab-eval.version: 0.5.1
instructlab-quantize.version: 0.1.0
instructlab-schema.version: 0.4.2
instructlab-sdg.version: 0.7.0
instructlab-training.version: 0.7.0
Torch:
torch.version: 2.3.1+cu121
torch.backends.cpu.capability: AVX2
torch.version.cuda: 12.1
torch.version.hip: None
torch.cuda.available: True
torch.backends.cuda.is_built: True
torch.backends.mps.is_built: False
torch.backends.mps.is_available: False
torch.cuda.bf16: True
torch.cuda.current.device: 0
torch.cuda.0.name: NVIDIA GeForce RTX 3070 Ti
torch.cuda.0.free: 7.1 GB
torch.cuda.0.total: 7.7 GB
torch.cuda.0.capability: 8.6 (see https://developer.nvidia.com/cuda-gpus#compute)
llama_cpp_python:
llama_cpp_python.version: 0.3.6
llama_cpp_python.supports_gpu_offload: False
Editing Generated Config
Use the command below to edit the generated config.
$ ilab config edit
Here we go through the file and ensure that we are setting the number of GPUs correctly. We also modify some additonal settings that have differed from the default (as ilab thinks we now have a $40k GPU). Below is some of what we have configured.
# Metadata pertaining to the specifics of the system which the Configuration is
# meant to be applied to.
metadata:
# Manufacturer, Family, and SKU of the system CPU, ex: Apple M3 Max
# Default: None
cpu_info:
# Amount of GPUs on the system, ex: 8
# Default: None
gpu_count: 1
# Family of the system GPU, ex: H100
# Default: None
gpu_family: RTX 3070 TI
# Manufacturer of the system GPU, ex: Nvidia
# Default: None
gpu_manufacturer: Nvidia
# Specific SKU related information about the given GPU, ex: PCIe, NVL
# Default: None
gpu_sku:
Working with Models
Before we start working with a model, lets take a look at the ilab command parameters. We can see that we have several options such as download, list, test, train, chat, serve, etc.

Lets download a model with the command below. The command below will download the default model sets in the config file
$ ilab model download
output below

View your downloaded models via the command below
$ ilab model list
Now we serve the default model with the command below
$ ilab model serve
Open another terminal window and source the env file from the instructlab directory
$ source venv/bin/activate
Then run the command below to begin an interactive chat session
$ ilab model chat
References
- https://developer.nvidia.com/cuda-downloads?target_os=Linux&target_arch=x86_64&Distribution=Ubuntu&target_version=22.04&target_type=deb_local
- https://github.com/instructlab/instructlab/issues/2422
- https://github.com/instructlab/instructlab
- https://developers.redhat.com/blog/2024/06/12/getting-started-instructlab-generative-ai-model-tuning#getting_started_with_instructlab
- https://www.techtarget.com/searchcio/definition/synthetic-data#:~:text=Synthetic%20data%20is%20increasingly%20used,fresh%20domain%20knowledge%20and%20explainability.
- https://docs.nvidia.com/cuda/cuda-installation-guide-linux/#meta-packages
- https://docs.redhat.com/en/documentation/red_hat_enterprise_linux_ai/1.2/html/building_your_rhel_ai_environment/initializing_instructlab#initialize_ilab

Leave a Reply