--- base_model: google/gemma-4-26B-A4B-it license: gemma language: en pipeline_tag: text-generation library_name: transformers tags: - text-generation - code - code-generation - secure-coding - vulnerability-detection - security - code-review - static-analysis - qlora - unsloth - gemma4 - conversational datasets: - bigcode/self-oss-instruct-sc2-exec-filter-50k --- # gemma-coder — security-hardened coding model `gemma-coder` is a QLoRA fine-tune of [`google/gemma-4-26B-A4B-it`](https://huggingface.co/google/gemma-4-26B-A4B-it) (a mixture-of-experts base with ~4B active parameters) that serves as the default agent model for the `remote-agent-dev-platform`. It runs inside a sandboxed agent that builds, executes, and tests everything it writes. The base model already writes competent code across the six languages this agent uses. **Correctness is no longer the frontier — safety is.** So the model's objective has been re-scoped from "write working code" to **write code that does not introduce vulnerabilities, and recognise the ones already in a codebase.** The measuring stick for that objective is [**ExploitGym**](https://rdi.berkeley.edu/blog/exploitgym). ## Why ExploitGym ExploitGym (Berkeley RDI, with the Max Planck Institute for Security and Privacy, UC Santa Barbara, Arizona State, Anthropic, OpenAI, and Google; arXiv [2605.11086](https://arxiv.org/abs/2605.11086)) is a benchmark of **898 real-world vulnerabilities** across three domains: | Domain | Instances | Source | | --- | --- | --- | | Userspace C/C++ (FFmpeg, OpenSSL, …) | 520 | OSS-Fuzz / OSV | | V8 JavaScript engine (Chromium) | 185 | real V8 bugs | | Linux kernel | 193 | privilege-escalation tasks | Each task ships vulnerable source, a proof-of-vulnerability input, and a containerised runtime with toggleable mitigations (ASLR, stack canaries, the V8 sandbox, KASLR). ExploitGym measures whether an agent can turn a *known* bug into a *working* exploit — the hardest possible probe of whether a model truly understands a vulnerability rather than pattern-matching its surface. We use it **defensively**. Every ExploitGym instance is a labelled, executable example of a real defect and exactly what it takes to trigger it. That is the richest possible curriculum for the two things this model is being aligned to do: 1. **Not emit those defect classes in the first place** — the exact memory-safety, type-confusion, and privilege-boundary mistakes ExploitGym is built from become negative examples the model learns to avoid as it writes. 2. **Find them in existing code** — given a diff or a file, locate the vulnerability, name its class (CWE), explain the trigger condition, and propose a fix, evaluated against ExploitGym's ground-truth bugs and its agent-as-a-judge verifier. The benchmark is used strictly as an **evaluation and alignment target for defensive security** (secure generation, vulnerability detection, and code review) — not to produce exploits. ExploitGym's own runs used structured, gated security-research access; this model does not reproduce that. ## Intended uses - **Secure code generation** across Python, JavaScript/React, Go, Java, and Swift, inside a sandbox that executes and tests every output. - **Vulnerability detection and code review** — flagging insecure patterns, naming the CWE class, and proposing fixes as it or a human writes. **Not intended for:** producing exploits or offensive tooling; safety-critical systems; or running generated code unreviewed. This is a small, free-tier-trained model — treat every output as a draft a human must review and test. ## Status & roadmap - **Shipped today:** the six-language coding QLoRA described under *Training* below, with the coding evaluation shown. - **In progress:** re-scoping toward the secure-coding + vulnerability-detection objective above, evaluated against ExploitGym. **Security-hardened results are not yet established — this card states the objective, not a completed benchmark.** It will be updated with ExploitGym-measured numbers once they exist; until then, no ExploitGym score should be attributed to this model. ## Training - **Method:** QLoRA (Unsloth), 4-bit, LoRA rank 8 / alpha 16 on attention layers, lr 2e-5, max seq len 768, AdamW-8bit. - **Compute:** weekly 8-hour sessions on Kaggle's free dual-T4 GPUs, resumed across cycles. - **Coding data:** [`bigcode/self-oss-instruct-sc2-exec-filter-50k`](https://huggingface.co/datasets/bigcode/self-oss-instruct-sc2-exec-filter-50k). ## Evaluation Current coding evaluation — multi-language sandboxed `pass@1` over executable test suites (promotion threshold 80%): | Language | pass@1 | | --- | --- | | Go | 4/4 | | Java | 4/4 | | JavaScript | 4/4 | | Python | 4/4 | | Swift | skipped (no toolchain in the eval image) | This is a small smoke-test suite (16 executed problems), not a broad coding benchmark — read it as a promotion gate, not a capability claim. The security objective above is evaluated separately against ExploitGym and is not yet reported. ## Limitations - Small, free-tier-trained MoE fine-tune; it can produce incorrect or insecure code. Always review and test before use. - The security re-scoping is in progress; do not rely on this model for vulnerability detection until ExploitGym-measured results are published here. - ExploitGym is referenced as a defensive evaluation target only. Do not use this model to generate exploits.