Instructions to use Aurumdev95/gemma4-31b-cessna-dpo-nvfp4 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Aurumdev95/gemma4-31b-cessna-dpo-nvfp4 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="Aurumdev95/gemma4-31b-cessna-dpo-nvfp4") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("Aurumdev95/gemma4-31b-cessna-dpo-nvfp4") model = AutoModelForMultimodalLM.from_pretrained("Aurumdev95/gemma4-31b-cessna-dpo-nvfp4", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use Aurumdev95/gemma4-31b-cessna-dpo-nvfp4 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Aurumdev95/gemma4-31b-cessna-dpo-nvfp4" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Aurumdev95/gemma4-31b-cessna-dpo-nvfp4", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/Aurumdev95/gemma4-31b-cessna-dpo-nvfp4
- SGLang
How to use Aurumdev95/gemma4-31b-cessna-dpo-nvfp4 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "Aurumdev95/gemma4-31b-cessna-dpo-nvfp4" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Aurumdev95/gemma4-31b-cessna-dpo-nvfp4", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "Aurumdev95/gemma4-31b-cessna-dpo-nvfp4" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Aurumdev95/gemma4-31b-cessna-dpo-nvfp4", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Unsloth Desktop
- Docker Model Runner
How to use Aurumdev95/gemma4-31b-cessna-dpo-nvfp4 with Docker Model Runner:
docker model run hf.co/Aurumdev95/gemma4-31b-cessna-dpo-nvfp4
gemma4-31b-cessna-dpo
A Gemma 4 (31B) assistant for the Cessna 172S, aligned with DPO to answer
in a clean, professional, pilot-friendly style. It is the preference-tuned successor
to the supervised model Aurumdev95/gemma4-31b-cessna,
whose knowledge was distilled from the Cessna 172S Pilot's Operating Handbook (POH).
⚠️ Study aid only. This model is for flight-training and familiarization. It is not an authoritative reference and must not be used for real-world flight operations, planning, or decision-making. Always consult the official POH and current approved documentation.
What DPO changed
The supervised model already answered POH questions accurately. DPO was applied to fix presentation: correct technical spelling, adopt consistent terminology, and format answers as well-structured Markdown (short bold labels, ordered steps, bullet lists) so they render cleanly and read the way a briefing would.
- Preference data: 133 train / 15 validation pairs. For each question the SFT model's own answer was the rejected response; an expert-editor rewrite that preserved every technical fact while improving clarity and formatting was the chosen response.
- Objective: DPO (
sigmoid) with anrpo_alpha = 1.0SFT anchor to prevent chosen-logprob collapse; 3 epochs,beta = 0.1, LoRAr = 16.
Results
Held-out preference accuracy rose each epoch — the model learned to prefer the professional, well-formatted answers:
| Epoch | Reward accuracy | Reward margin |
|---|---|---|
| 1 | 75.0% | 0.16 |
| 2 | 81.3% | 0.38 |
| 3 | 87.5% | 0.41 |
Crucially, factual accuracy did not regress. On the same 46-question held-out POH validation set (served in NVFP4 via vLLM, greedy decoding):
| Metric | SFT | DPO |
|---|---|---|
| token-F1 (mean) | 0.478 | 0.488 |
| ROUGE-L (mean) | 0.374 | 0.385 |
| answers with F1 ≥ 0.3 | 42/46 | 41/46 |
The lexical-overlap scores are essentially unchanged (the DPO answers are longer and more structured, which these terse-reference metrics do not reward), while the style is markedly more professional — the intended outcome.
Variants
| Repo | Format | Use |
|---|---|---|
Aurumdev95/gemma4-31b-cessna-dpo |
LoRA adapter (this repo) | merge onto the SFT model |
Aurumdev95/gemma4-31b-cessna-dpo-nvfp4 |
NVFP4 (Blackwell) | vLLM on RTX 50 / GB10 |
Aurumdev95/gemma4-31b-cessna-dpo-fp8 |
FP8 (W8A8) | vLLM, broad GPU support |
Prompting
Use a flight-assistant system prompt, e.g.:
You are a knowledgeable flight assistant for the Cessna 172S. Answer questions about the aircraft's systems, limitations, and procedures accurately and concisely, based on the Pilot's Operating Handbook.
Training stack
Unsloth + TRL 0.24 + Transformers 5.5 on an NVIDIA DGX Spark (GB10). See the training pipeline for full configuration.
- Downloads last month
- 20
Model tree for Aurumdev95/gemma4-31b-cessna-dpo-nvfp4
Base model
google/gemma-4-31B