# QLoRA weight-budget experiment

This is a CPU-only arithmetic experiment, not a trainer or a GPU-memory benchmark. Python standard library only; tested with Python 3.11.5 on macOS arm64. No model, dataset, dependency installation or network access is needed.

Keep `qlora_memory.py` and `test_qlora_memory.py` together. Run:

```sh
python3 -B -m unittest -v test_qlora_memory.py
python3 -B qlora_memory.py --parameters 7000000000 --rank 16 --projection 4096:4096:64 --budget-gib 8
```

`--projection IN:OUT:COUNT` groups distinct, unshared linear weight matrices with identical dimensions. Repeat the flag to add another group. For the article's seven-projection illustration, use `--projection 4096:4096:128 --projection 4096:11008:96`. Set `--rank 32` for the third row. These input shapes are illustrative, not a checkpoint's verified configuration.

The base payload assumes all supplied base parameters occupy four bits, rounded up to whole bytes. LoRA adds `rank * (in_features + out_features)` parameters per target matrix. The default adapter storage is four bytes per parameter; use `--adapter-bytes 2`, `4` or `8` to state a different assumption. Compute dtype is not adapter storage dtype. Base and adapter weights are assumed resident together, without offload or sharding.

The program rejects impossible nonpositive dimensions/counts, missing target matrices, and a sum of targeted base weights greater than the supplied total base parameter count. It cannot inspect a model to detect duplicate targets or verify your dimensions.

Excluded: quantization metadata, extra bytes for weights retained above four bits, gradients, optimizer states, activations, temporary buffers, allocator/runtime overhead. Bias tuning, embedding adapters, shared adapters, variable ranks, DoRA and other adapter variants are not modeled.

An under-budget weight floor returns **undetermined**, never a successful training-fit prediction. An over-budget floor rules out keeping the assumed representations resident together within that budget. A real recipe still needs GPU measurement and a separate model-quality evaluation.

Units: one GB is 1,000,000,000 bytes; one GiB is 1,073,741,824 bytes. The budget flag accepts a positive finite decimal GiB value, rounded down to whole bytes. Exit 0 means a valid calculation, including an over-budget result; invalid inputs exit 2.
