--- library_name: transformers pipeline_tag: text-generation license: apache-2.0 base_model: - Qwen/Qwen3.6-27B base_model_relation: finetune language: - en - zh tags: - veriloop - veriloop-coder - code - coding-agent - software-engineering - repository-understanding - tool-use - peft - lora - safetensors - self-harness - harness-engineering - surface-host-adapter - evidence-binding - rollback - uncertainty-calibration - long-context - open-source - apache-2.0 - vertical-code-model - recursive-improvement ---
System Boundary · Runtime Protocol · Benchmarks · Self-Harness Contract · Surface Host · Loading Toolkit · Licensing
Evaluation metadata · Evaluation evidence repository · Full-resolution figure
Figure 1. Comparative benchmark snapshot as of 27 July 2026. Panel (a) restricts the comparison to models below 32B parameters; panel (b) shows the all-model view represented in the published snapshot. The figure is a static comparative record and may not reflect subsequent leaderboard updates.
These are **agent-system results**, not checkpoint-only measurements. The prototype runtime generated the task artifacts, while each benchmark's native evaluator determined the reported outcome. Where published, the evidence packages bind task identity, generated artifact, execution or evaluation record, and integrity metadata to support task-level inspection. Publication of a score, figure, or evidence package does not by itself imply endorsement, certification, or independent verification by a benchmark maintainer. ## Voder: Development Direction **Voder** is the working name of the general-purpose coding agent under development by the VeriLoop team. It is intended to integrate the private Self-Harness control plane with VeriLoop Coder-E1 so that repository state, tool execution, evidence admission, validation, targeted repair, rollback, candidate selection, and long-running task state are governed as one auditable software-engineering process. The public repository provides the model-side foundation for that direction. The four benchmark evaluations are early internal systems experiments in Voder's technical lineage; they are not claims that the open checkpoint reproduces the complete agent, nor that the evaluated prototype is the final Voder product. The long-term research objective is **bounded recursive improvement**: validated failure and repair evidence should improve how later tasks are framed, investigated, checked, and corrected. This is not unrestricted self-modification. The recursion remains constrained by current-task supremacy, evidence admission, explicit budgets, deterministic validation, rollback, safe stopping, benchmark-locked evaluation, and one selected deliverable with a verifiable evidence chain. Voder remains under active development. Its eventual release scope, availability, capabilities, commercial terms, and schedule have not yet been determined. ## Public Self-Harness Functional Contract The following describes the **public functional contract** of the VeriLoop Self-Harness. It specifies the required engineering stages and evidence semantics, but not the private Voder implementation, prompts, thresholds, routing code, or arbitration policy. ```text Current task ↓ Contract compilation ↓ Evidence admission → Candidate realization ↓ Falsification → Gap-driven exploration ↓ Surgical repair → Re-verification ↓ Selection or rollback ↓ One final deliverable + next-loop evidence ``` ### Phase A — Contract compilation Resolve the task family, repository scope, required artifact, interfaces, constraints, acceptance conditions, uncertainty, tool budget, and explicit halt conditions into one active engineering contract. ### Phase B — Evidence admission and candidate realization Admit only repository context, tests, traces, tool receipts, selected references, and prior prevention rules that can change the artifact. Generate a complete candidate under that bounded evidence state. ### Phase C — Falsification and counterevidence Challenge the candidate against task intent, repository interfaces, tests, contradictions, unsupported assumptions, and observable failure signals. Syntactic validity or plausibility alone is not sufficient. ### Phase D — Gap-driven exploration Trigger retrieval, search, reverse analysis, sandbox execution, or additional model work only for a decisive unresolved gap. Every action must have trigger evidence, an expected output, and a defined downstream consumer. ### Phase E — Surgical repair and re-verification Convert the observed failure into a focused repair constraint, correct the smallest broken invariant, preserve unaffected surfaces, and re-run the same contract-aware checks. ### Phase F — Selection, rollback, and evidence inheritance Select exactly one eligible deliverable, or stop and roll back when the evidence does not support delivery. Persist validation receipts, failure class, repair outcome, halt reason, and reusable prevention evidence for later loops. The operating semantics are compact: ```text Evidence → Falsification → Exploration → Repair ↺ ``` Re-verification completes Repair; a verified outcome becomes next-loop Evidence. This is how VeriLoop converts generation into bounded, auditable engineering rather than uncontrolled retry. --- ## Expert Convergence Mode **Expert Mode** uses a richer bounded convergence loop for difficult software-engineering tasks. It is designed for: - repository-scale changes with hidden invariants; - ambiguous failures requiring evidence synthesis; - cross-file API or behavior changes; - tasks with multiple plausible repairs; - cases where a first candidate should be challenged before delivery. At a functional level, Expert Mode adds: - richer contract and evidence compilation; - multiple candidate opportunities where justified; - independent validation and failure classification; - focused repair cycles; - candidate comparison and quality arbitration; - strict selected-artifact delivery; - explicit safe stopping when no candidate satisfies the required threshold. The exact prompts, thresholds, orchestration order, candidate scoring, repair policies, and routing logic are proprietary. ### Runtime Modes | Mode | Optimization target | Functional behavior | |---|---|---| | **Flash** | Lowest latency | Direct artifact realization with lightweight control and minimal orchestration | | **Thinking** | Balanced quality and cost | Evidence-guided artifact generation with bounded self-check and targeted correction | | **Expert** | Highest final quality | Rich evidence, iterative validation and repair, candidate arbitration, strict selected delivery | --- ## Surface Host Adapter The **Surface Host Adapter** is VeriLoop's detachable model-side adaptation plane. It hosts the published PEFT artifacts as explicit behavioral control surfaces without redefining or mathematically rewriting the frozen backbone. The four published PEFT surfaces are model-side behavioral control artifacts designed to interoperate with the private Self-Harness interfaces used in the Voder development lineage. They specialize signal formation for tool contracts, uncertainty, rollback, and evidence binding; they do not contain or reproduce the orchestration, routing, validation, repair, arbitration, or stopping policy of the private runtime. This layer must be understood in the correct architectural order: ```text Self-Harness = core governing technology Surface Host Adapter = external adaptation host Narrow-domain PEFT artifacts = hosted behavioral surfaces ``` The PEFT surfaces are therefore **not** the core of VeriLoop, not a substitute for the backbone, and not a substitute for the Self-Harness. They are purpose-built auxiliary components whose value is realized when the Harness uses their signals to improve contract sensitivity, evidence handling, escalation, repair, rollback, and delivery control. ### External injection model The public PEFT artifacts remain logically separate from the distributed model weights. The current published adapters target explicit `surface_host.*` modules rather than the raw backbone's attention or MLP projections. They must therefore be hosted through the public Surface Host scaffold and wrapper; they are not ordinary backbone LoRA adapters and must not be presented as if they were mathematically merged into the Qwen checkpoint. The original backbone remains unchanged. The optional packaging utility performs a mergeability preflight. For the current `surface_host.*` adapters, its correct result is a **co-packaged deployment bundle** containing the backbone assets and four detachable adapter directories—not a mathematical weight merge. This architecture provides three important properties: - **separation of concerns** — the backbone retains general coding and reasoning capability, while each PEFT surface contributes a narrowly defined behavioral bias; - **detachable composition** — published surfaces can be loaded, removed, evaluated, or replaced independently of the original checkpoint; - **Harness-centered control** — the Self-Harness remains the governing layer that determines how task contracts, evidence, validation, and repair consume the resulting surface signals. The Surface Host Adapter is not a single fifth adapter. It is the host boundary for multiple specialized adaptation surfaces. This design separates three kinds of capability: - **backbone capability** — code understanding, generation, reasoning, and language competence; - **surface capability** — narrow-domain sensitivity learned through PEFT; - **Harness capability** — evidence admission, falsification, exploration, repair, re-verification, arbitration, and recursive evidence evolution. The current public adaptation surfaces are: | Surface | Primary optimization | Practical effect | |---|---|---| | **ToolSpec** | Tool schemas, argument constraints, preconditions, postconditions, execution-facing formats | Stronger sensitivity to malformed calls, missing prerequisites, invalid arguments, and incomplete execution contracts | | **Uncertainty** | Answer, evidence, execution, specification, and risk uncertainty | Better escalation decisions, reduced unsupported certainty, and more selective use of search, tools, or validation | | **Rollback** | Validator negation, failed edits, bounded correction, state restoration | More precise repair behavior and lower risk of broad destructive rewrites after a local failure | | **Evidence Binding** | Claim-to-evidence alignment, provenance, validation context, source discipline | Stronger coupling between generated artifacts, supporting context, and observable verification evidence | ### What the adaptation plane improves The narrow-domain tuning is intended to materially strengthen: - tool-call and schema discipline; - contract adherence before generation; - repository-aware code planning; - evidence-sensitive reasoning; - uncertainty-triggered escalation; - validator-aware repair behavior; - rollback and correction boundaries; - artifact-only delivery discipline; - claim, source, and validation alignment; - consistency across long, multi-stage coding workflows. These adapters are optimized as **behavioral control surfaces**, not as isolated leaderboard specialists. Their main value appears when they are hosted by the Surface Host Adapter and coordinated by the Self-Harness, where adapter signals can influence routing, evidence injection, repair control, and delivery policy. Adapter-only loading does not reproduce the complete production system. --- ## Public Rule Lineage: Karpathy → Forrest Chang → Mnilax → Libo Wang The public VeriLoop rule system has a documented intellectual lineage. 1. **Andrej Karpathy's original observations.** In January 2026, Karpathy publicly described recurring coding-agent failure modes: silent assumptions, unmanaged confusion, overcomplication, unnecessary adjacent edits, and weak success criteria. 2. **Forrest Chang's four-rule operationalization.** Forrest Chang converted those observations into a compact `CLAUDE.md` behavior contract containing four principles: **Think Before Coding, Simplicity First, Surgical Changes, and Goal-Driven Execution**. These are referred to here as the Karpathy-origin **Golden Four**, while credit for packaging them into the formal four-rule repository belongs to Forrest Chang. 3. **Mnilax's eight agent-era additions.** In May 2026, Mnilax published eight additional controls for newer agentic failure modes: - use the model for judgment calls, and keep deterministic decisions in code; - enforce hard token and execution budgets; - surface conflicting patterns instead of averaging them; - read relevant code before writing; - make tests verify intent rather than appearance; - checkpoint significant multi-step work; - follow repository conventions unless explicitly changing them; - fail visibly rather than silently reporting success. 4. **VeriLoop's adaptation.** VeriLoop materially rewrites and extends this 4+8 lineage into a **14-rule public Self-Harness contract** organized around task supremacy, evidence admission, falsification, targeted exploration, surgical repair, deterministic enforcement, checkpointing, traceability, and domain-overlay isolation. The VeriLoop rules are not presented as Karpathy's, Forrest Chang's, or Mnilax's exact text, nor as an official collaboration with those authors. They are an attributed adaptation for VeriLoop's evidence-bound Self-Harness architecture. ### Primary references - Andrej Karpathy's original X post: https://x.com/karpathy/status/2015883857489522876 - Forrest Chang's four-rule repository (currently hosted under `multica-ai`): https://github.com/multica-ai/andrej-karpathy-skills - Mnilax's original eight-rule extension on X, published 9 May 2026: https://x.com/Mnilax/status/2053116311132155938 ## The Libo Wang 14-Rule Public Self-Harness Contract The following rules are intentionally public. They form the model-visible engineering discipline shared across VeriLoop coding modes. 1. **Current-Task Supremacy** The current request and its exact output contract override old memory, cached templates, prior habits, and unrelated retrieved context. 2. **Escalate on Evidence, Not Instinct** Trigger search, reverse analysis, sandbox execution, or repair only when concrete uncertainty, missing evidence, or a failed contract justifies the cost. 3. **Inspect Before Editing** Read the relevant entry points, callers, interfaces, tests, repository conventions, selected evidence, and failure signals before changing code. 4. **Produce the Minimal Complete Artifact** Deliver the smallest implementation that fully satisfies the task, preserves required interfaces, and can be validated. 5. **Repair the Broken Invariant, Not the Whole System** Prefer a precise correction of the failing region. Rewrite broadly only when evidence proves that local repair cannot restore correctness. 6. **Validate Intent, Not Appearance** Syntax, formatting, and imports are necessary checks; the decisive test is whether the artifact satisfies the user's actual functional intent. 7. **Surface Failure; Never Simulate Success** Keep failures, skipped checks, degraded states, and unknowns explicit in Harness evidence. Never claim execution, validation, or success that did not occur. 8. **Put Determinism in Code** Parsing, static checks, validation, scoring, self-tests, and reproducible transformations belong in typed deterministic code whenever possible. 9. **Treat Budget as an Execution Contract** Use token, tool, time, and compute budgets deliberately. The requested deliverable receives priority over commentary, duplicated context, and optional explanation. 10. **Make Every Tool Call Accountable** Every tool, search, retrieval, or execution action must have trigger evidence, an expected output, and a defined downstream consumer. 11. **Checkpoint Long-Running Work** Persist candidate artifacts, validation receipts, repair records, selection results, and progress state so useful work survives interruption and remains auditable. 12. **Preserve Local and Task-Family Conventions** Respect filenames, APIs, paths, repository style, artifact format, language conventions, benchmark constraints, and user-defined operating rules. 13. **Deliver One Artifact with a Verifiable Evidence Chain** Select exactly one final deliverable while retaining the evidence bundle that explains why it was chosen. 14. **Separate Core Discipline from Domain Overlays** Apply specialized domain or benchmark rules only when the current task requires them, and never allow an overlay to override the current request. --- ## Core Technical Advantages ### Evidence binding over prompt accumulation VeriLoop does not equate more context with better context. Raw findings are compiled into constraints, selected evidence, and validation targets before they reach generation. ### Native artifact routing Repository patches, polyglot source files, shell tasks, functions, configuration files, and structured outputs are handled as distinct artifact families rather than being forced into a single Python-centric path. ### Deterministic–generative separation The model handles ambiguity, synthesis, and repair hypotheses. Deterministic components handle parsing, structural checks, contract enforcement, receipts, and reproducible transformations. ### Failure-aware convergence Validation failures are preserved as evidence. The system distinguishes an invalid candidate, a missing dependency, a degraded tool, an unverified assumption, and a genuine task failure instead of collapsing them into a generic retry. ### Safe stopping When evidence does not support delivery, the system can stop with an explicit failure record rather than manufacturing a confident result. ### Task-level traceability Evaluation packages can bind a task identity to the generated artifact and the corresponding official evaluation record, allowing third parties to inspect the complete task-level chain. --- ## Current Public Repository Contents The repository contains the open model component of the VeriLoop system, published adaptation artifacts, and selected evaluation and traceability materials. Unless a file states otherwise, the repository's original VeriLoop artifacts are distributed under the Apache License 2.0, subject to applicable third-party licenses and notices. ### Backbone and runtime-compatible files - sharded `safetensors` model weights; - model and generation configuration; - tokenizer and preprocessing assets; - Hugging Face-compatible loading metadata; - public-safe loading, hosting, verification, and packaging utilities: - `load_veriloop_coder_e1_surface_host_adapter_publicsafe.py`; - `veriloop_coder_peft_host_publicsafe.py`; - `surface_host_wrapper_publicsafe.py`; - `merge_veriloop_adapters_into_backbone_publicsafe.py`. ### Narrow-domain adaptation artifacts Each public Surface Host PEFT directory may include: - adapter weights and adapter configuration; - tokenizer assets; - best-checkpoint records; - epoch history; - host manifests; - adapter plans; - training-result summaries; - sanitized training and evaluation records; - training manifests; - specialized surface heads where published. Public Surface Host PEFT roots: ```text surface_host_peft_toolspec_adapter/adapter/ surface_host_peft_uncertainty_adapter/adapter/ surface_host_peft_rollback_adapter/adapter/ surface_host_peft_evidence_binding_adapter/adapter/ ``` These directories contain the published narrow-domain fine-tuning products used as external Surface Host Adapter attachments. They were developed to align with Voder's private task-contract, evidence, uncertainty, rollback, and verification interfaces. They are support layers for Self-Harness operation, not standalone substitutes for either the Self-Harness or Voder. ### Evaluation and traceability artifacts The repository also publishes selected evidence packages for: - SWE-bench Verified; - SWE-bench Pro; - Terminal-Bench 2.0; - DeepSWE. These packages are designed to support task-level inspection across: ```text task identity → prototype-generated artifact → native evaluation record ``` Publication of an evidence package does not imply independent third-party verification unless explicitly stated. --- ## What Remains Private The Apache 2.0 release applies to the model weights and public repository artifacts. It does not automatically license private VeriLoop components that are not distributed here. For strategic and competitive reasons, the production Self-Harness and the developing Voder runtime remain private at this stage. The protected boundary includes: - exact Self-Harness orchestration and stage implementation; - Voder runtime code, service interfaces, deployment packages, and operator tooling; - internal prompt compiler and model-visible packet construction; - evidence routing, ranking, admission, and contamination controls; - tool scheduling, repository-state management, and execution governance; - candidate scoring, repair arbitration, rollback policy, and delivery thresholds; - private memory schemas, evidence inheritance, and evolution policies; - sandbox, worktree, permission, and observability infrastructure; - full proprietary training data and data-construction pipelines; - internal benchmark routing and anti-overfitting controls. The public model card describes the functional architecture and exposes bounded implementation guidance; it does not release the executable production Harness. The public 14-rule contract is a behavioral interface disclosure, not the Voder runtime. Use of the public model and published artifacts follows their stated Apache 2.0 terms. Access to non-public Voder or production Self-Harness components, if offered, requires a separate written agreement with the applicable rights holder or authorized project entity. --- ## Model Overview | Property | Value | |---|---| | Model family | VeriLoop Coder-E1 | | System product direction | Voder — general-purpose coding agent under development | | Developed by | Tsinghua SIGS Robot Lab | | Technical Lead & AI Researcher | Libo Wang | | Contact | [free.equality.anyone@gmail.com](mailto:free.equality.anyone@gmail.com) | | Backbone | Qwen3.6-27B | | Parameter class | 27B backbone (<32B); additional PEFT-based Surface Host Adapters | | Release status | Open-source model distribution under Apache License 2.0 | | System boundary | Open Coder-E1 weights and public artifacts; private production Self-Harness and Voder runtime | | Core technical contribution | Self-Harness for governed evidence-loop execution within the VeriLoop/Voder system | | Adaptation | External Surface Host Adapter with narrow-domain PEFT/control surfaces | | Primary domain | Software engineering, coding agents, repository repair, tool-mediated code generation | | Languages | English, Chinese | | Weight format | `safetensors` | | Runtime modes | Flash, Thinking, Expert | | Output families | Patches, source files, functions, scripts, configurations, structured answers | | Benchmark attribution | Internal prototype agent runtime + pre-release Self-Harness + VeriLoop Coder-E1 + applicable Surface Host PEFT + benchmark-native execution/evaluation stack | | Evaluation philosophy | Task-level evidence, official evaluation records, traceability packages | --- ## Recommended Use Cases VeriLoop Coder-E1 is intended for: - repository understanding and codebase navigation; - bug localization and surgical repair; - patch drafting and validation; - cross-file API changes; - tool-mediated software-engineering agents; - terminal and automation tasks; - validator-aware repair workflows; - evidence-grounded coding assistance; - benchmark and evaluation research; - long-running engineering tasks requiring checkpoints and auditability. --- ## Public Loading & Surface Host Toolkit > **Public scope.** These utilities expose the released model, the four published Surface Host PEFT artifacts, and their public loading boundary. > > **Private boundary.** They do not include Voder, the production Self-Harness, orchestration policies, routing thresholds, prompt compiler, repair arbitration, benchmark runtime, or commercial serving stack. > > **Runtime baseline.** Because the private Self-Harness is not included, first preserve the repository-native protocol by using the designated Qwen3.6-27B revision, repository-shipped tokenizer, and built-in `chat_template.jinja`, with matching parsers and EOS/stop semantics. Once this baseline is verified, developers may integrate a third-party client or build an independent Harness using the guidance in [Runtime Protocol and Recovery Guidance](#runtime-protocol-and-recovery-guidance). ### Toolkit at a glance | Utility | Primary role | Main output | |---|---|---| | `load_veriloop_coder_e1_surface_host_adapter_publicsafe.py` | Validate one published adapter without loading the full 27B backbone | Adapter verification manifest | | `veriloop_coder_peft_host_publicsafe.py` | Load the frozen backbone once and instantiate the explicit public `surface_host` scaffold | PEFT host manifest | | `surface_host_wrapper_publicsafe.py` | Load all four published `surface_host.*` adapter weights and run a bounded host-side forward smoke test | Surface Host report and signal packet | | `merge_veriloop_adapters_into_backbone_publicsafe.py` | Perform mergeability preflight and build a deployable model-and-adapter package | Deployment package and package report | ### Repository layout Run the commands from the repository root after downloading the model files and all four published adapter directories: ```text . ├── config.json ├── model.safetensors.index.json ├── veriloop-coder-e1-model-*.safetensors ├── load_veriloop_coder_e1_surface_host_adapter_publicsafe.py ├── veriloop_coder_peft_host_publicsafe.py ├── surface_host_wrapper_publicsafe.py ├── merge_veriloop_adapters_into_backbone_publicsafe.py ├── surface_host_peft_toolspec_adapter/ │ └── adapter/ ├── surface_host_peft_uncertainty_adapter/ │ └── adapter/ ├── surface_host_peft_rollback_adapter/ │ └── adapter/ └── surface_host_peft_evidence_binding_adapter/ └── adapter/ ``` ### Environment preparation ```python python -m pip install torch transformers peft safetensors accelerate huggingface_hub mkdir -p public_loading_reports ``` --- ### Step 1 — Validate one published adapter **Purpose:** verify the adapter configuration, Safetensors files, tensor metadata, base-model binding, and public security boundary without loading the complete 27B backbone. **Example: Evidence Binding** ```python python load_veriloop_coder_e1_surface_host_adapter_publicsafe.py \ --base-model . \ --adapter-source . \ --adapter-subfolder surface_host_peft_evidence_binding_adapter/adapter \ --verify-only \ --manifest-out public_loading_reports/evidence_binding_adapter_manifest.json ``` **Output** ```text public_loading_reports/evidence_binding_adapter_manifest.json ``` Repeat the same command with any of the following adapter subfolders: ```text surface_host_peft_toolspec_adapter/adapter surface_host_peft_uncertainty_adapter/adapter surface_host_peft_rollback_adapter/adapter surface_host_peft_evidence_binding_adapter/adapter ``` > **Technical boundary:** the current adapters target explicit `surface_host.*` modules. They are not ordinary raw-backbone LoRA targets and must not be presented as direct `PeftModel.from_pretrained(base_model, ...)` attachments to the unmodified Qwen backbone. --- ### Step 2 — Load the frozen backbone and create the public PEFT host **Purpose:** load the distributed backbone exactly once, preserve the frozen-backbone boundary, instantiate the explicit `surface_host` module tree, and export a public host manifest. ```python python veriloop_coder_peft_host_publicsafe.py \ --backbone . \ --local-files-only \ --dtype bf16 \ --device-map auto \ --manifest-out public_loading_reports/peft_host_manifest.json ``` **Output** ```text public_loading_reports/peft_host_manifest.json ``` > **Security default:** remote model code is disabled. Enable it only for a pinned repository revision that genuinely requires custom code and has been independently reviewed. --- ### Step 3 — Load and execute all four Surface Host adapter weights **Purpose:** load the ToolSpec, Uncertainty, Rollback, and Evidence Binding Safetensors into detachable host-side modules and run a bounded forward smoke test. ```python python surface_host_wrapper_publicsafe.py \ --repo-root . \ --runtime-mode full \ --timeout-s 120 \ --emit-report public_loading_reports/surface_host_report.json \ --emit-injection-text public_loading_reports/surface_host_signal_packet.txt ``` **Outputs** ```text public_loading_reports/surface_host_report.json public_loading_reports/surface_host_signal_packet.txt ``` This command: - loads all four published adapter families independently; - verifies real adapter-weight hosting and host-side forward execution; - keeps the raw backbone unchanged; - does not monkey-patch or rename backbone modules; - does not disclose the private production routing policy that consumes the resulting signals. --- ### Step 4 — Build a deployable model-and-adapter package **Purpose:** perform mergeability preflight and create one deployment directory containing the model assets and all four detachable Surface Host adapters. ```python python merge_veriloop_adapters_into_backbone_publicsafe.py \ --base-model . \ --toolspec-adapter surface_host_peft_toolspec_adapter/adapter \ --evidence-adapter surface_host_peft_evidence_binding_adapter/adapter \ --rollback-adapter surface_host_peft_rollback_adapter/adapter \ --uncertainty-adapter surface_host_peft_uncertainty_adapter/adapter \ --output-dir veriloop_coder_e1_surface_host_package \ --preflight-mergeability \ --package-if-not-mergeable \ --base-file-mode hardlink_or_copy \ --adapter-file-mode copy \ --force \ --backup-existing \ --json-out public_loading_reports/surface_host_package_report.json ``` **Expected result for the current release** ```text packaged_not_mathematically_merged ``` The current adapters target `surface_host.*`, not raw-backbone attention or MLP projections. The technically correct deployment artifact is therefore a package containing the frozen backbone assets and four detachable adapter directories—not a claim that the adapters were mathematically merged into the original checkpoint. ### Recommended public workflow ```text Adapter verification (optional) │ ▼ Frozen-backbone / host compatibility check │ ▼ Four-adapter Surface Host loading and forward smoke test │ ▼ Deployment packaging (optional) ``` > These scripts expose only the released model-and-adapter boundary. They do not reproduce Voder. Developers may use the public architecture, 14-rule contract, and technical essay to design an independent Harness, but equivalence with the private Voder runtime or its benchmark behavior is not guaranteed. ## Limitations - The distributed backbone, the four Surface Host PEFT artifacts, and the public utilities do not reproduce Voder or the production Self-Harness. - The reported benchmark scores are system-level results from an internal prototype agent configuration and must not be interpreted as checkpoint-only performance or as results from a final Voder product. - Independent Harness implementations may differ materially in routing, tool use, evidence quality, repair policy, stopping behavior, and final results. - Community-modified chat templates, mismatched parsers, custom EOS/stop rules, or third-party client behavior can cause premature termination, malformed tool calls, or duplicated final messages even when the model and adapters are loaded correctly. - Without an independent validation, rollback, and bounded-retry loop, recoverable protocol faults may be exposed directly to the user because the production Self-Harness is not included in the public release. - Coding outputs may still be incorrect, incomplete, insecure, or incompatible with the target environment. - Validation quality depends on available repository context, tests, tools, permissions, and execution environments. - Long-context operation requires appropriate accelerator memory and KV-cache planning. - External documentation and retrieved evidence may be stale, incomplete, or conflicting. - Safe stopping reduces unsupported delivery but cannot eliminate all false positives or false negatives. - Published evaluation evidence should be interpreted according to its stated verification status. --- ## Safety and Responsible Use Generated code should be treated as an engineering proposal until validated. Recommended safeguards: - run generated code in isolated environments; - inspect shell commands and dependency changes before execution; - use tests, static analysis, security review, and repository-specific checks; - preserve rollback points for destructive operations; - require human review for high-impact or security-sensitive changes; - do not infer successful execution from plausible-looking output; - keep credentials, private repositories, and sensitive logs outside uncontrolled prompts. --- ## Open-Source Release and Licensing **VeriLoop Coder-E1 is released as an open-source model under the Apache License 2.0.** The model weights and original VeriLoop artifacts distributed in this repository may be used, reproduced, modified, redistributed, and deployed—including in commercial applications—provided that users comply with the Apache License 2.0 and preserve any required license and notice information. This grant covers the public repository artifacts, including: - distributed model weights; - model configuration and tokenizer assets; - published Surface Host PEFT adapter artifacts; - public-safe loading, hosting, verification, and packaging utilities; - original manifests, compatibility files, and published evaluation materials, except where a file states a different license. Third-party components, upstream model materials, datasets, benchmark assets, and externally sourced content remain subject to their respective licenses and terms. The Apache 2.0 release does **not** include Voder, the private production Self-Harness implementation, internal prompts, orchestration code, evidence-routing logic, repair and rollback policy, private training data, proprietary data-construction pipelines, or serving infrastructure unless those components are published separately under an explicit license. ```text Included under Apache-2.0: public Coder-E1 weights + published Surface Host PEFT + public repository artifacts Not included in this release: private Self-Harness + Voder runtime + unpublished internal systems ``` Access to or deployment of non-public Voder and production Self-Harness components, if offered, requires a separate written agreement with the applicable rights holder or authorized project entity. This requirement does not reduce or override the Apache 2.0 rights attached to the public files actually distributed in this repository. Use of the open model or its published PEFT surfaces does not imply endorsement, certification, warranty, or support by VeriLoop Lab, Tsinghua University, or any affiliated organization. The public artifacts are provided under the terms and warranty limitations of the Apache License 2.0. --- ## Citation ```bibtex @misc{veriloop_coder_e1_2026, title = {VeriLoop Coder-E1: Evidence-Bound Self-Harness Loops for Recursive Software Engineering}, author = {Wang, Libo}, year = {2026}, note = {Developed by Tsinghua SIGS Robot Lab; Technical Lead and AI Researcher: Libo Wang}, howpublished = {Hugging Face model repository}, url = {https://huggingface.co/tsinghua-sigs-robot-lab/veriloop-coder-e1} } ``` --- ## Acknowledgements VeriLoop Coder-E1 builds on a Qwen3.6-27B-compatible foundation and the broader open machine-learning tooling ecosystem. The public Harness discipline acknowledges: - **Andrej Karpathy**, whose public observations identified recurring coding-agent failure modes; - **Forrest Chang**, who operationalized those observations into the compact four-principle `andrej-karpathy-skills` repository; - **Mnilax**, who published eight additional agent-era rules in May 2026; - the communities behind Transformers, PEFT, Safetensors, vLLM, software-engineering benchmarks, repository-level evaluation, and reproducible model deployment. VeriLoop's 14-rule contract is a materially adapted public technology layer of the Self-Harness. The production Self-Harness and Voder runtime—including proprietary orchestration, prompts, routing, scoring, repair, memory, training systems, and production controls—remain private. --- ## A Note from the Author VeriLoop was built through an unconventional research path. I entered this work without a conventional software-engineering background, and the project developed through skepticism, rejection, repeated failure, and sustained iteration. The lesson I take from that experience is not that persistence can replace rigor. It is that original hypotheses, disciplined evidence, transparent limitations, and engineering execution can allow unconventional researchers to produce work that deserves evaluation on its technical merits. I present VeriLoop neither as a claim to personal infallibility nor as a request for approval. I present it as an argument for intellectual independence: criticism should be answered with reproducible artifacts, explicit boundaries, stronger evidence, and better systems. — **Libo Wang**