OneRuby.devAN ENGINEERING NOTEBOOK

AI · 6 min read

QLoRA on a Single GPU: Does a 7B Model Fit in 8GB?

Use a tested Python memory calculator to separate four-bit weights from QLoRA training memory, and measure the parts an 8GB claim leaves out.

Can you fine-tune a 7B model with QLoRA on a single 8GB GPU? The parameter count and the phrase “4-bit” cannot answer that. They describe part of the weight storage. A training step also needs memory for work that happens around those weights.

This note starts with the part we can calculate, using a small Python program that runs on a CPU. It counts an ideal four-bit base payload and ordinary LoRA adapter weights. It does not predict peak VRAM or establish that a training recipe fits. No GPU training run was performed.

The 3.5GB number has a narrow meaning

For exactly 7,000,000,000 parameters, four bits per parameter gives 3,500,000,000 bytes: 3.5GB, or about 3.260GiB. GB means a billion bytes; GiB means 1,073,741,824 bytes. Keep the unit when comparing a calculation with a device report.

That calculation assumes every base parameter occupies exactly four bits. It excludes quantization metadata and any weights kept at a higher precision. It also says nothing about the memory needed to use those weights.

QLoRA's original paper describes a frozen, quantized base model with trainable low-rank adapters. NF4 represents base weights; double quantization reduces the storage of quantization constants; paged optimizers address memory spikes. Quantized storage and computation are different concerns. The paper's results belong to its reported experiments; they are not an 8GB guarantee for an arbitrary model and dataset.

Count adapters from shapes

For an ordinary LoRA adapter on an out_features × in_features weight matrix, the two added matrices contain rank × (in_features + out_features) parameters. Repeat that for each distinct target matrix. Bias training, embedding adapters, shared adapters and other variants need different accounting.

The calculator uses this function:

Python
def lora_parameter_count(projection, rank):
require_positive_integer("rank", rank)
return projection.count * rank * (
projection.in_features + projection.out_features
)

A 4096 × 4096 projection at rank 16 adds 131,072 parameters. But “rank 16” does not specify how many projections receive adapters, or their shapes. A target with a narrower output dimension adds fewer parameters. Count the actual modules selected in your model; do not assume every attention projection is square.

For a controlled example, take an illustrative seven-billion-parameter model with 32 blocks. First, adapt two 4096 × 4096 projections per block: 64 matrices in total. These are calculator inputs, not a downloaded checkpoint's configuration.

Save qlora_memory.py and test_qlora_memory.py in the same directory. They use Python's standard library, without PyTorch or a model download. Run:

Terminal
python3 -B qlora_memory.py --parameters 7000000000 --rank 16 --projection 4096:4096:64 --budget-gib 8

The projection argument means input dimension:output dimension:number of matrices. Here the budget is eight GiB of assumed capacity; use the available bytes reported by your device when doing a real comparison. The observed output was:

Output
Ideal 4-bit base payload: 3,500,000,000 bytes (3.260 GiB)
LoRA parameters: 8,388,608
Adapter weights (4 bytes each): 33,554,432 bytes (0.031 GiB)
Weight-storage floor: 3,533,554,432 bytes (3.291 GiB)
Excluded: quantization metadata, unquantized-weight uplift, gradients,
optimizer states, activations, temporary buffers, allocator/runtime overhead.
Budget verdict: undetermined: training overhead is not estimated

The four-byte adapter assumption is explicit and adjustable with --adapter-bytes. It is not inferred from a compute dtype. Inspect the actual adapter storage dtype before using the result in a budget.

Now target seven projections per block: four 4096 × 4096 matrices and three rectangular matrices with dimensions 4096 and 11008. For adapter counting, swapping the input and output dimensions leaves their sum unchanged. The corresponding arguments are --projection 4096:4096:128 --projection 4096:11008:96.

These results came from running the calculator, with the same base count and four bytes per adapter parameter:

Targets in the illustrative modelRankAdapter parametersAdapter weight bytesCombined weight floor
Two square projections per block168,388,60833,554,4323.291GiB
Seven projections per block1639,976,960159,907,8403.409GiB
Seven projections per block3279,953,920319,815,6803.557GiB

Doubling rank doubles the adapter count here. It leaves the base payload unchanged. None of these rows measures a training step, and none establishes which adapter choice gives better model quality.

A lower bound can reject a budget, but cannot approve it

The calculator deliberately returns undetermined when the weight floor is below the supplied budget. The gap is unaccounted memory, not measured spare capacity. With a 3GiB budget, the first example instead reports that the weight floor already exceeds it.

Both conclusions assume the supplied base and adapter weights reside together, without offloading or sharding. They use the storage representations supplied to the calculation. This is a budget check under stated assumptions, not a universal hardware requirement.

The test suite checks exact counts, rectangular projections, rank and storage changes, byte rounding, and the budget boundary. It rejects invalid dimensions, nonpositive ranks, nonfinite budgets and target matrices whose base-weight count exceeds the supplied model total. A dedicated test prevents a small weight floor from becoming a “training fits” verdict.

Terminal
python3 -B -m unittest -v test_qlora_memory.py

The recorded CPU run used Python 3.11.5: 18 tests passed. The example notes list the assumptions and commands.

Keep the training API tied to a version

The earlier recipe on this page combined trl>=0.21.0 with older trainer keywords. The TRL 0.21.0 reference places max_length and packing in SFTConfig; the tokenizer or processor is passed to SFTTrainer as processing_class. That version also documents eval_strategy. An open-ended installation constraint does not preserve a copied example's API.

PEFT 0.17.0's quantization guide documents preparing a quantized model with prepare_model_for_kbit_training before adding LoRA. Its LoraConfig reference defines what target_modules="all-linear" selects, including the output-layer exclusion for a PreTrainedModel. Check those matches before counting adapters.

These are versioned documentation checks, not a tested training environment. This article supplies no GPU training script or package compatibility claim.

What to measure on the GPU you intend to use

To turn “fits in 8GB” into a reproducible result, record the model revision, quantization settings, adapter targets and dtypes, package versions, GPU, driver and CUDA environment. Then measure the actual recipe:

  1. Record loading separately from training. After loading, keep that phase's peak before resetting the counters for training.
  2. Run complete forward, backward and optimizer steps. Include the first optimizer step and several later steps, using representative long batches as well as typical ones. Test evaluation and saving separately if the workflow needs them.
  3. Record peak tensor allocation and peak allocator reservation. PyTorch's max_memory_allocated reports the former; reset_peak_memory_stats sets the start of the measurement interval. Reserved memory includes memory managed by the caching allocator. These are different measurements, not two quantities to add together.
  4. Record device-level usage too. PyTorch's profiler cannot see allocations made outside its allocator. A tensor peak alone is not total GPU usage.
  5. Change one setting at a time: sequence length, microbatch size, target modules, rank, checkpointing or attention implementation. Keep the input and measurement interval with each result.

This is a proposed measurement procedure; it was not executed here. A successful run would establish one tested configuration. Until that evidence exists, the honest answer for a particular 7B model on an 8GB GPU remains open. The calculator helps locate what is known before the training experiment begins.

Found a mistake or tried a different approach?

Send Alex a note ↗