Llama-3.2-1B QLoRA Summarizer (SAMSum)

A merged fine-tune of Llama-3.2-1B trained with QLoRA (4-bit) on the SAMSum dialogue summarization dataset. This is the full merged model (base weights + LoRA adapter), so it can be loaded and used directly without peft.

Model Details

  • Base model: meta-llama/Llama-3.2-1B
  • Fine-tuning method: QLoRA (4-bit NF4 quantized training, adapter merged into base weights post-training)
  • Task: Dialogue-to-summary generation
  • Dataset: SAMSum (~14K dialogue/summary pairs)

Best hyperparameters (selected via systematic HPT):

Hyperparameter Value
LoRA rank (r) 16
LoRA alpha 32
Learning rate 2e-4
Target modules q_proj, k_proj, v_proj, o_proj
Training steps 300 (~1 epoch)
Batch size 4 (grad accumulation = 4)

Results

Model ROUGE-1 ROUGE-2 ROUGE-L
Llama-3.2-1B (base) 34.18% 11.52% 27.24%
This model 47.41% 24.06% 39.55%

How to Use

from transformers import AutoModelForCausalLM, AutoTokenizer
import torch

model_id = "MohamedQiqa/Llama-3.2-1B-QLoRA-Summarizer"

tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
    model_id,
    torch_dtype=torch.bfloat16,
    device_map="auto",
)

dialogue = """
Amanda: I baked cookies. Do you want some?
Jerry: Sure! Bring them over.
Amanda: I'll be there in 10 minutes.
"""

messages = [
    {"role": "system", "content": "Summarize the following dialogue in one or two sentences."},
    {"role": "user", "content": dialogue},
]

inputs = tokenizer.apply_chat_template(
    messages, add_generation_prompt=True, return_tensors="pt"
).to(model.device)

output = model.generate(inputs, max_new_tokens=64, do_sample=False)
summary = tokenizer.decode(output[0][inputs.shape[-1]:], skip_special_tokens=True)
print(summary)

Quantized inference (optional, lower memory)

from transformers import BitsAndBytesConfig

bnb_config = BitsAndBytesConfig(
    load_in_4bit=True,
    bnb_4bit_quant_type="nf4",
    bnb_4bit_compute_dtype=torch.bfloat16,
)

model = AutoModelForCausalLM.from_pretrained(
    model_id,
    quantization_config=bnb_config,
    device_map="auto",
)

Intended Use

  • Short dialogue/conversation summarization (chat logs, meeting notes, messenger exchanges)
  • English text only
  • Best suited for short-to-medium length dialogues similar in style to SAMSum

Limitations

  • Trained on a single epoch over SAMSum; may not generalize well to very long or technical documents
  • Inherits any biases present in the base Llama-3.2-1B model and the SAMSum dataset
  • Not instruction-tuned for tasks beyond summarization

Training Details

Fine-tuned using QLoRA (4-bit quantization + LoRA adapters) on a single GPU, with hyperparameters selected through a one-at-a-time sensitivity sweep over LoRA rank, learning rate, and target modules, tracked in Weights & Biases.

We optimized 1B model now outperforms the base GPT-4o-mini

Model Type ROUGE-1 ROUGE-2 ROUGE-L
Llama 3.2 1B (base) Open-weight baseline 34.18% 11.52% 27.24%
GPT-4o-mini (base) Frontier baseline 40.76% 16.15% 32.91%
Llama 3.2 1B (Lesson 4) First fine-tune (r8, q+v) 46.74% 23.17% 38.88%
Llama 3.2 1B (optimized) Best HPT config 🏆 47.41% 24.06% 39.55%
GPT-4o-mini (fine-tuned) Managed fine-tuned 53.68% 30.73% 45.96%

Experiment 1: How Much Capacity Do Adapters Need?

image

Experiment 2: How Fast Should the Model Learn?

image

Experiment 3: Which Model Components Should We Train?

Target Modules ROUGE-1 ROUGE-2 ROUGE-L Δ vs Baseline
q + v 46.74% 23.17% 38.88% baseline
Full attention 47.41% 24.06% 39.55% +0.67% ✅🏆
Attention + MLP 47.62% 23.73% 39.23% +0.35%

image

The Winning Configuration

Putting it all together, the best setup for Llama 3.2 1B on SAMSum is:

Hyperparameters:

  • LoRA rank: 16
  • LoRA alpha: 32
  • earning rate: 2e-4
  • Target modules: q, k, v, o
  • Training steps: 300 (~1 epoch)
  • Batch size: 4 (grad accumulation = 4)
Downloads last month
13
Safetensors
Model size
1B params
Tensor type
F32
·
U8
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for MohamedQiqa/Llama-3.2-1B-QLoRA-Summarizer

Adapter
(751)
this model