Summarization
Transformers
Safetensors
English
llama
text-generation
lora
qlora
fine-tuned
dialogue-summarization
text-generation-inference
4-bit precision
bitsandbytes
Instructions to use MohamedQiqa/Llama-3.2-1B-QLoRA-Summarizer with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use MohamedQiqa/Llama-3.2-1B-QLoRA-Summarizer with Transformers:
# Use a pipeline as a high-level helper # Warning: Pipeline type "summarization" is no longer supported in transformers v5. # You must load the model directly (see below) or downgrade to v4.x with: # 'pip install "transformers<5.0.0' from transformers import pipeline pipe = pipeline("summarization", model="MohamedQiqa/Llama-3.2-1B-QLoRA-Summarizer")# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("MohamedQiqa/Llama-3.2-1B-QLoRA-Summarizer") model = AutoModelForCausalLM.from_pretrained("MohamedQiqa/Llama-3.2-1B-QLoRA-Summarizer", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Llama-3.2-1B QLoRA Summarizer (SAMSum)
A merged fine-tune of Llama-3.2-1B trained with QLoRA (4-bit) on the SAMSum dialogue summarization dataset. This is the full merged model (base weights + LoRA adapter), so it can be loaded and used directly without peft.
Model Details
- Base model:
meta-llama/Llama-3.2-1B - Fine-tuning method: QLoRA (4-bit NF4 quantized training, adapter merged into base weights post-training)
- Task: Dialogue-to-summary generation
- Dataset: SAMSum (~14K dialogue/summary pairs)
Best hyperparameters (selected via systematic HPT):
| Hyperparameter | Value |
|---|---|
| LoRA rank (r) | 16 |
| LoRA alpha | 32 |
| Learning rate | 2e-4 |
| Target modules | q_proj, k_proj, v_proj, o_proj |
| Training steps | 300 (~1 epoch) |
| Batch size | 4 (grad accumulation = 4) |
Results
| Model | ROUGE-1 | ROUGE-2 | ROUGE-L |
|---|---|---|---|
| Llama-3.2-1B (base) | 34.18% | 11.52% | 27.24% |
| This model | 47.41% | 24.06% | 39.55% |
How to Use
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch
model_id = "MohamedQiqa/Llama-3.2-1B-QLoRA-Summarizer"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
model_id,
torch_dtype=torch.bfloat16,
device_map="auto",
)
dialogue = """
Amanda: I baked cookies. Do you want some?
Jerry: Sure! Bring them over.
Amanda: I'll be there in 10 minutes.
"""
messages = [
{"role": "system", "content": "Summarize the following dialogue in one or two sentences."},
{"role": "user", "content": dialogue},
]
inputs = tokenizer.apply_chat_template(
messages, add_generation_prompt=True, return_tensors="pt"
).to(model.device)
output = model.generate(inputs, max_new_tokens=64, do_sample=False)
summary = tokenizer.decode(output[0][inputs.shape[-1]:], skip_special_tokens=True)
print(summary)
Quantized inference (optional, lower memory)
from transformers import BitsAndBytesConfig
bnb_config = BitsAndBytesConfig(
load_in_4bit=True,
bnb_4bit_quant_type="nf4",
bnb_4bit_compute_dtype=torch.bfloat16,
)
model = AutoModelForCausalLM.from_pretrained(
model_id,
quantization_config=bnb_config,
device_map="auto",
)
Intended Use
- Short dialogue/conversation summarization (chat logs, meeting notes, messenger exchanges)
- English text only
- Best suited for short-to-medium length dialogues similar in style to SAMSum
Limitations
- Trained on a single epoch over SAMSum; may not generalize well to very long or technical documents
- Inherits any biases present in the base Llama-3.2-1B model and the SAMSum dataset
- Not instruction-tuned for tasks beyond summarization
Training Details
Fine-tuned using QLoRA (4-bit quantization + LoRA adapters) on a single GPU, with hyperparameters selected through a one-at-a-time sensitivity sweep over LoRA rank, learning rate, and target modules, tracked in Weights & Biases.
We optimized 1B model now outperforms the base GPT-4o-mini
Model Type ROUGE-1 ROUGE-2 ROUGE-L Llama 3.2 1B (base) Open-weight baseline 34.18% 11.52% 27.24% GPT-4o-mini (base) Frontier baseline 40.76% 16.15% 32.91% Llama 3.2 1B (Lesson 4) First fine-tune (r8, q+v) 46.74% 23.17% 38.88% Llama 3.2 1B (optimized) Best HPT config 🏆 47.41% 24.06% 39.55% GPT-4o-mini (fine-tuned) Managed fine-tuned 53.68% 30.73% 45.96%
Experiment 1: How Much Capacity Do Adapters Need?
Experiment 2: How Fast Should the Model Learn?
Experiment 3: Which Model Components Should We Train?
| Target Modules | ROUGE-1 | ROUGE-2 | ROUGE-L | Δ vs Baseline |
|---|---|---|---|---|
| q + v | 46.74% | 23.17% | 38.88% | baseline |
| Full attention | 47.41% | 24.06% | 39.55% | +0.67% ✅🏆 |
| Attention + MLP | 47.62% | 23.73% | 39.23% | +0.35% |
The Winning Configuration
Putting it all together, the best setup for Llama 3.2 1B on SAMSum is:
Hyperparameters:
- LoRA rank: 16
- LoRA alpha: 32
- earning rate: 2e-4
- Target modules: q, k, v, o
- Training steps: 300 (~1 epoch)
- Batch size: 4 (grad accumulation = 4)
- Downloads last month
- 13
Model tree for MohamedQiqa/Llama-3.2-1B-QLoRA-Summarizer
Base model
meta-llama/Llama-3.2-1B

