Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
82 commits
Select commit Hold shift + click to select a range
85f2047
test update to lumi latest container
mahnoormahnoorr Jul 3, 2026
18db59a
added use_cache= False for better results
mahnoormahnoorr Jul 6, 2026
09f54d2
updated script to new lumi aif container
mahnoormahnoorr Jul 6, 2026
e882be7
updated script to latest lumi aif container
mahnoormahnoorr Jul 6, 2026
6989c4b
Update README.md
mahnoormahnoorr Jul 6, 2026
47435cc
Update README.md
mahnoormahnoorr Jul 6, 2026
cf0f9b5
minor changes
mahnoormahnoorr Jul 6, 2026
52ffec1
minor changes
mahnoormahnoorr Jul 6, 2026
9de4ced
Create run-bnb-quantization-roihu.sh
mahnoormahnoorr Jul 17, 2026
6718b8c
Rename run-bob-quantization-roihu.sh to run-bnb-quantization-roihu.sh
mahnoormahnoorr Jul 17, 2026
d3ac0cd
Create run-gptq-config-roihu.sh
mahnoormahnoorr Jul 17, 2026
a438e2d
Update README.md
mahnoormahnoorr Jul 17, 2026
09770fc
Update README.md
mahnoormahnoorr Jul 17, 2026
69673ea
Create gptq-config-roihu.py
mahnoormahnoorr Jul 17, 2026
958a791
Create run-gptq-modifier-roihu.sh
mahnoormahnoorr Jul 17, 2026
6c83449
Update README.md
mahnoormahnoorr Jul 17, 2026
b77bf0d
Update README.md
mahnoormahnoorr Jul 17, 2026
20f7910
Create run-awq-modifier-roihu.sh
mahnoormahnoorr Jul 17, 2026
3f69b92
Update run-gptq-config-roihu.sh
mahnoormahnoorr Jul 17, 2026
a359378
Update run-gptq-modifier-roihu.sh
mahnoormahnoorr Jul 17, 2026
8f1b71b
Update README.md
mahnoormahnoorr Jul 17, 2026
04f6d5b
Update README.md
mahnoormahnoorr Jul 17, 2026
8e3df10
Update README.md
mahnoormahnoorr Jul 17, 2026
11e0043
Update README.md
mahnoormahnoorr Jul 17, 2026
e016eb1
Update README.md
mahnoormahnoorr Jul 17, 2026
3e23c35
Update README.md
mahnoormahnoorr Jul 17, 2026
bb2f7b4
Update README.md
mahnoormahnoorr Jul 17, 2026
bd40b09
Update README.md
mahnoormahnoorr Jul 17, 2026
48c3783
Update README.md
mahnoormahnoorr Jul 17, 2026
1284a49
Update README.md
mahnoormahnoorr Jul 17, 2026
680959d
Update README.md
mahnoormahnoorr Jul 17, 2026
298185b
Update README.md
mahnoormahnoorr Jul 17, 2026
1db80cc
Update README.md
mahnoormahnoorr Jul 17, 2026
c28f474
Update README.md
mahnoormahnoorr Jul 17, 2026
d577020
Update README.md
mahnoormahnoorr Jul 17, 2026
064b759
Update run-gptq-config-roihu.sh
anni-moisala Jul 20, 2026
0fc2524
Delete GPTQ/gptq-config.py
mahnoormahnoorr Jul 20, 2026
214f7a8
Rename gptq-config-roihu.py to gptq-config.py
mahnoormahnoorr Jul 20, 2026
d543c4b
Update run-gptq-config-roihu.sh
mahnoormahnoorr Jul 20, 2026
9103d3a
Update run-bnb-quantization-roihu.sh
mahnoormahnoorr Jul 20, 2026
9a6638e
Update run-gptq-modifier-roihu.sh
mahnoormahnoorr Jul 20, 2026
6957f5c
Update run-awq-modifier-roihu.sh
mahnoormahnoorr Jul 20, 2026
63f63e5
Update README.md
mahnoormahnoorr Jul 20, 2026
7fa2ecb
Update README.md
mahnoormahnoorr Jul 20, 2026
d51b908
remove puhti and mahti scripts
anni-moisala Jul 21, 2026
1587377
remove remaining puhti and mahti scripts
anni-moisala Jul 21, 2026
396d836
Update README.md
mahnoormahnoorr Jul 23, 2026
b20e7fd
Update README.md
mahnoormahnoorr Jul 23, 2026
b91ec83
Update README.md
mahnoormahnoorr Jul 23, 2026
15cf172
Update README.md
mahnoormahnoorr Jul 23, 2026
5004d3f
Update README.md
mahnoormahnoorr Jul 24, 2026
1c593fa
Update README.md
mahnoormahnoorr Jul 24, 2026
a041cb1
Update README.md
mahnoormahnoorr Jul 24, 2026
092c1e8
Update README.md
mahnoormahnoorr Jul 24, 2026
f0d201e
Update README.md
mahnoormahnoorr Jul 24, 2026
0eccd4a
Update README.md
mahnoormahnoorr Jul 27, 2026
f771bfc
Update README.md
mahnoormahnoorr Jul 27, 2026
42d64be
Update README.md
anni-moisala Jul 29, 2026
b5b3d1b
Update README.md
anni-moisala Jul 29, 2026
7520965
fix formatting
anni-moisala Jul 29, 2026
6d75884
Update README.md
anni-moisala Jul 29, 2026
f5c84f4
small fixes
anni-moisala Jul 29, 2026
47983f7
Update README.md
anni-moisala Jul 29, 2026
e51b8b9
updated script to latest released container
mahnoormahnoorr Aug 10, 2026
ace163f
minor update
mahnoormahnoorr Aug 10, 2026
bfa1cef
minor update
mahnoormahnoorr Aug 10, 2026
c8120d9
Update README.md
mahnoormahnoorr Aug 10, 2026
fbeff7f
updates container
mahnoormahnoorr Aug 10, 2026
c7e6337
updates to container: lumi-multitorch-u24r70f21m50t210-20260731_122833
mahnoormahnoorr Aug 10, 2026
f5eea14
minor update
mahnoormahnoorr Aug 10, 2026
4c29e29
minor update
mahnoormahnoorr Aug 10, 2026
ac2d0a0
change the model_name
mahnoormahnoorr Aug 11, 2026
20fcad4
major changes in the script
mahnoormahnoorr Aug 11, 2026
895adf8
major updates
mahnoormahnoorr Aug 11, 2026
26d8b32
minor changes
mahnoormahnoorr Aug 11, 2026
ba105a4
minor update
mahnoormahnoorr Aug 11, 2026
56ef3f4
minor update
mahnoormahnoorr Aug 11, 2026
e1695af
updated script to remove venv
mahnoormahnoorr Aug 11, 2026
c7504bc
updated
mahnoormahnoorr Aug 11, 2026
0ee1d0e
Update README.md
anni-moisala Aug 12, 2026
4ed35d5
Update README.md
anni-moisala Aug 12, 2026
b2095f5
Update README.md
anni-moisala Aug 12, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
38 changes: 28 additions & 10 deletions AWQ/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -6,33 +6,51 @@ In order to target weight and activation scaling locations within the model, the

## Installations

## LUMI

To run AWQ quantization scripts on LUMI, you can use the `gptqmodel` library that already ships inside the LUMI AI Factory container — no extra packages or virtual environment are required for AWQ. `torch` , `transformers`, and `datasets` are also already provided in the container.

The script loads the `Singularity` container environment and sets the container image path:

```bash
module purge
module use /appl/local/laifs/modules
module load lumi-aif-singularity-bindings

export SIF=/appl/local/laifs/containers/lumi-multitorch-u24r70f21m50t210-20260731_122833/lumi-multitorch-full-u24r70f21m50t210-20260731_122833.sif

```

---

## Roihu

The CSC preinstalled PyTorch module covers most of the libraries needed to run these examples
(torch, transformers, datasets, accelerate). The rest can be installed on top of the module in a virtual environment.
(torch, transformers, datasets, accelerate). Llmcompressor can be installed on top of the module in a virtual environment.

### Load the module
```bash
module purge
module use /appl/local/csc/modulefiles
module load pytorch/2.7
module load python-pytorch/2.10
```
### Create and activate a virtual environment using system packages
```bash
python3 -m venv --system-site-packages venv
source venv/bin/activate
```
### Install packages

```bash
pip install optimum==1.27.0 llmcompressor==0.7.1 --cache-dir ./.pip-cache
```
The flag --cache-dir points the pip cache to the current (scratch) folder instead of the default (home directory), to avoid filling up home directory quota.
(venv)> pip install llmcompressor==0.12.0 --cache-dir ./.pip-cache
```
---

## Usage

The launch scripts are:

- `run-awq-modifier-lumi.sh` - quantizes model on LUMI with 1 GPU
- `run-awq-modifier-mahti.sh` - quantizes model on Mahti with 1 GPU
- `run-awq-modifier-puhti.sh` - quantizes model on Puhti with 1 GPU
- `run-awq-modifier-lumi.sh` - quantizes model on LUMI with 1 GPU
- `run-awq-modifier-roihu.sh` - quantizes model on Roihu with 1 GPU

**Note:** the scripts are made to be run on `gputest` or `dev-g` partition with a 30 minutes time-limit. You have to select the proper partition for longer jobs for your real runs. Additionally, change the `--account` parameter to your own project code.

Expand All @@ -55,6 +73,6 @@ Meaning the script quantizes the model’s linear layers using a mixed-precision
- Model size (MB) before and after quantization.

## Notes
- The current scripts use **Falcon-RW-1B** for fast experimentation. You can replace `model_name` with a larger model. In this case, you might want to disable saving the full model.
- The current scripts use **TinyLlama-1.1B-Chat-v1.0** for fast experimentation. You can replace `model_name` with a larger model. In this case, you might want to disable saving the full model.
- For large models, `device_map="auto"` allows the model modules to be moved between the CPU and GPU for quantization.
- Feel free to experiment with different values for `num_calibration_samples` and `max_seq_lenght` and to modify the quantization recipe.
2 changes: 1 addition & 1 deletion AWQ/awq-modifier.py
Original file line number Diff line number Diff line change
Expand Up @@ -11,7 +11,7 @@
from llmcompressor.modifiers.awq import AWQModifier
from llmcompressor.utils import dispatch_for_generation

model_name = "tiiuae/falcon-rw-1b"
model_name = "TinyLlama/TinyLlama-1.1B-Chat-v1.0"
dataset_name = "HuggingFaceH4/ultrachat_200k"
dataset_split = "train_sft"
num_calibration_samples = 256
Expand Down
11 changes: 6 additions & 5 deletions AWQ/run-awq-modifier-lumi.sh
Original file line number Diff line number Diff line change
Expand Up @@ -11,15 +11,16 @@

# Load the module
module purge
module use /appl/local/csc/modulefiles
module load pytorch/2.7
module use /appl/local/laifs/modules
module load lumi-aif-singularity-bindings

# Activate the virtual environment from your current directory or change to the appropriate path
source venv/bin/activate
# export path to used container image
export SIF=/appl/local/laifs/containers/lumi-multitorch-u24r70f21m50t210-20260731_122833/lumi-multitorch-full-u24r70f21m50>

# Activate the virtual environment from your current directory or change to th
# This will store all the Hugging Face cache such as downloaded models
# and datasets in the project's scratch folder
export HF_HOME=/scratch/${SLURM_JOB_ACCOUNT}/${USER}/hf-cache
mkdir -p $HF_HOME

srun python3 awq-modifier.py
srun singularity exec "$SIF" python3 awq-modifier.py
26 changes: 0 additions & 26 deletions AWQ/run-awq-modifier-mahti.sh

This file was deleted.

26 changes: 0 additions & 26 deletions AWQ/run-awq-modifier-puhti.sh

This file was deleted.

22 changes: 22 additions & 0 deletions AWQ/run-awq-modifier-roihu.sh
Original file line number Diff line number Diff line change
@@ -0,0 +1,22 @@
#!/bin/bash
#SBATCH --account=project_xxxxxxx
#SBATCH --partition=gputest
#SBATCH --time=00:15:00
#SBATCH --nodes=1
#SBATCH --tasks-per-node=1
#SBATCH --cpus-per-task=72
#SBATCH --gres=gpu:gh200:1
#SBATCH --mem=32G

module purge
module load python-pytorch/2.10

# Activate the virtual environment from your current directory or change to the appropriate path
source venv/bin/activate

# Set hf cache to the project's scratch
export HF_HOME=/scratch/$SLURM_JOB_ACCOUNT/$USER/hf-cache/
mkdir -p $HF_HOME

srun python3 awq-modifier.py

7 changes: 3 additions & 4 deletions BitsAndBytes/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -4,15 +4,14 @@ This example demonstrates quantizing the **OPT-125M** model using the [bitsandby

## Running the script

All of the libraries needed to run this example (transformers, bitsandbytes, accelerate) are covered by the CSC preinstalled PyTorch module.
All of the libraries needed to run this example (transformers, bitsandbytes, accelerate) are covered by the AI Factory provided Container on LUMI or the CSC preinstalled PyTorch module on Roihu.

The script `bnb-quantization.py` will quantize the OPT-125M model to nf4 or NormalFloat 4-bit, introduced to use with QLoRA technique, a parameter efficient fine-tuning technique. It can be used with QLoRA for fine-tuning, or without just for reducing model size.

The launch scripts are:

- `run-bnb-quantization-lumi.sh` - quantizes model on LUMI with 1 GPU
- `run-bnb-quantization-mahti.sh` - quantizes model on Mahti with 1 GPU
- `run-bnb-quantization-puhti.sh` - quantizes model on Puhti with 1 GPU
- `run-bnb-quantization-lumi.sh` - quantizes model on LUMI with 1 GPU
- `run-bnb-quantization-roihu.sh` - quantizes model on Roihu with 1 GPU

**Note:** the scripts are made to be run on `gputest` or `dev-g` partition with a 30 minutes time-limit. You have to select the proper partition for longer jobs for your real runs. Additionally, change the `--account` parameter to your own project code.

Expand Down
1 change: 1 addition & 0 deletions BitsAndBytes/bnb-quantization.py
Original file line number Diff line number Diff line change
Expand Up @@ -27,6 +27,7 @@ def benchmark(model, tokenizer, prompt):
**inputs,
max_new_tokens=50,
do_sample=True,
use_cache=False,
temperature=0.7,
)

Expand Down
12 changes: 8 additions & 4 deletions BitsAndBytes/run-bnb-quantization-lumi.sh
Original file line number Diff line number Diff line change
Expand Up @@ -11,12 +11,16 @@

# Load the module
module purge
module use /appl/local/csc/modulefiles
module load pytorch/2.7
module use /appl/local/laifs/modules
module load lumi-aif-singularity-bindings

# export path to used container image
export SIF=/appl/local/laifs/containers/lumi-multitorch-u24r70f21m50t210-20260731_122833/lumi-multitorch-full-u24r70f21m50t210-20260731_122833.sif

# This will store all the Hugging Face cache such as downloaded models
# and datasets in the project's scratch folder
export HF_HOME=/scratch/${SLURM_JOB_ACCOUNT}/${USER}/hf-cache
export HF_HOME=/scratch/$SLURM_JOB_ACCOUNT/$USER/llm-quantization-scripts/BitsAndBytes/hf-cache
mkdir -p $HF_HOME
export SINGULARITYENV_HF_HOME=$HF_HOME

srun python3 bnb-quantization.py
srun singularity exec "$SIF" bash -c 'python3 bnb-quantization.py'
23 changes: 0 additions & 23 deletions BitsAndBytes/run-bnb-quantization-mahti.sh

This file was deleted.

23 changes: 0 additions & 23 deletions BitsAndBytes/run-bnb-quantization-puhti.sh

This file was deleted.

21 changes: 21 additions & 0 deletions BitsAndBytes/run-bnb-quantization-roihu.sh
Original file line number Diff line number Diff line change
@@ -0,0 +1,21 @@
#!/bin/bash
#SBATCH --account=project_xxxxxxx
#SBATCH --partition=gputest
#SBATCH --time=00:15:00
#SBATCH --nodes=1
#SBATCH --tasks-per-node=1
#SBATCH --cpus-per-task=72
#SBATCH --gres=gpu:gh200:1
#SBATCH --mem=32G

module purge
module load python-pytorch/2.10

# Activate the virtual environment from your current directory or change to the appropriate path
source venv/bin/activate

# Set hf cache to the project's scratch
export HF_HOME=/scratch/$SLURM_JOB_ACCOUNT/$USER/hf-cache/
mkdir -p $HF_HOME

srun python3 bnb-quantization.py
Loading