Laxmi β experimental software-state prediction
Predict the consequences of a software change before executing it.
Created and maintained by Navneet Prabhakar, creator and sole contributor to Laxmi. Third-party libraries retain their authorship and licenses.
Laxmi investigates whether observed software state, history, an action and a proposed code change can predict what happens next. This release keeps the bounded controller experiment inspectable and adds local real-patch evaluation tooling to the same research project. It is not a chat model, a general code-review assistant or an autonomous software engineer.
Version 0.1.1 Β· Experimental research release
Controller weights and runtime are unchanged from v0.1.0. The latest model repair failed its improvement test. The new real-patch exact-execution tool is a baseline, not a trained model. None of the four controller candidates is promoted as an accepted model.
The goal
Laxmi's long-term ambition is a testable software world model: a system that maintains a predictive view of software state as evidence arrives. Given the current state, execution history, a proposed action or code change, and declared operating conditions, it should forecast possible consequences before execution, express uncertainty, and revise its view when new observations arrive. A useful forecast would help an engineer decide what to inspect, test, change or leave alone.
The research path has three stages:
- Predictive state: learn action- and source-conditioned outcomes, including bounded multi-step effects and differences between proposed changes. Measure forecasts against actual outcomes and strong methods given the same information.
- Prospective advice: save forecasts before real controller actions, then test whether the advice improves a human review decision at acceptable error, coverage, latency and cost.
- Broader engineering assessment: use source and pull-request context, observed history and explicit constraints to assess likely consequences, compare a finite set of alternatives, and explain findings with traceable evidence.
The ultimate aspiration is useful engineer-like work within a defined role and measurable constraints. Progress is measured by prospective outcomes and decisions, including honest abstention when evidence is insufficient. This v0.1.1 release is an early, bounded controller experiment toward the first stage. It has not established a general software world model, a qualified advisory, or autonomous engineering capability.
The experiment
A controller records progress and retries failed operations. An edit to its recovery logic can delay completion or repeat work. The research question is whether pre-action evidence can predict that consequence.
State + history + source/edit + declared conditions
β
Save a prediction
β
Execute a controlled experiment
β
Compare with observed outcomes
The released consumer forecasts admitted synthetic controller conditions. It can observe, predict, compare proposals, save/restore state and score separately recorded outcomes. It never dispatches native controller actions. Recording a forecast with this consumer alone does not authenticate prospective ordering.
Results, including the failures
| Seed | Original H1 exact outcomes | Normalized H1 exact outcomes | Changed pairs both correct, either readout |
|---|---|---|---|
| 17 | 8 / 24 | 8 / 24 | 0 / 6 |
| 43 | 4 / 24 | 4 / 24 | 0 / 6 |
Readout normalization did not improve the task. These are exposed development results from September 15β16, 2026, not independent acceptance scores. Longer-horizon likelihood comparisons remain unresolved. Probabilities are not established as calibrated, and independent implementation transfer remains open.
The packaged example checks installation and replay. It must not be presented as a fresh benchmark. See verification.json for the exact smoke scope and BENCHMARK.md for the next comparison design.
Real-patch development screen
This release also contains local real-patch evaluation tooling. Six exposed historical changes were inventoried; only one currently has a valid forecast packet, and five have oracle-only scenarios excluded from forecast accuracy. On the one scored development case, the same-information exact-execution baseline matched both branch observations and the patch effect. The scenario was selected after its outcome was known. There are no blind/acceptance-eligible cases, no learned real-patch model scored, and no released case data. The result shows that this fully runnable case offers no measured headroom over exact execution; it does not validate generalization.
Download and run
Controller runtime verified platform: Apple Silicon macOS, Python 3.12.11. That wheel is macOS ARM64 tagged. Other Python 3.12 patch builds and platforms have not been verified for the controller runtime. The source archive is provided for inspection and experimentation; it does not establish Linux controller support. No GPU is required for the example.
With the Hugging Face hf CLI installed:
hf download Executespec/laxmi-controller-experimental --revision v0.1.1 --local-dir laxmi-release
cd laxmi-release
shasum -a 256 -c SHA256SUMS
python3.12 -m venv .venv
.venv/bin/python -m pip install runtime/laxmi_readout_runtime_capsule-0.1.0-py3-none-macosx_11_0_arm64.whl
.venv/bin/python -I -B verify_release.py --release . --python "$PWD/.venv/bin/python" --output smoke-17-original --candidate 17-original
Expected completion: 17-original: all six operations passed. The output directory must be new. The check observes the example, predicts, compares, restores, saves and scores an outcome. Omit --candidate to check all four candidates. The verifier makes candidate/example files private (0600) as required by the runtime. It writes local snapshots and logs; it does not run the controlled native software or train a model.
The wheel declares exact dependency versions, also recorded in requirements-macos-arm64-py312.txt. Dependencies are installed separately; they are not bundled. Version pins are not a dependency artifact hash lock. The source package includes the runtime source and license files. No optimizer checkpoint is shipped.
The separate real-patch package has its own local installation instructions. It installed and validated its input in an offline Linux ARM64 container, but Linux exact execution and any hosted API remain unqualified. Its OpenAPI file is a design draft only.
Package layout
| Path | Contents |
|---|---|
candidates/17-original/, 17-normalized/, 43-original/, 43-normalized/ |
Original retained safetensors weights, configuration, manifests and matching consumer metadata |
runtime/ |
Reviewed runtime wheel and source ZIP |
evaluation/real-patch/ |
Local exact-baseline evaluator wheel, source, contract and design-only API draft; no model weights or case data |
examples/ |
One retained synthetic observation, proposal comparison and outcome template |
verify_release.py |
Installation/replay check using only this package |
release.json, SHA256SUMS |
Candidate identity, runtime identity and file integrity |
verification.json |
Sanitized installation/replay evidence |
LICENSE, NOTICE, LICENSE_SCOPE.md, RELEASE_SCOPE.md |
Apache 2.0, attribution and exact distribution scope |
CITATION.cff |
Suggested research citation |
Weights, configuration, readout metadata and runtime must match. Original and normalized readouts share tensor layouts but differ in computation. Do not mix them. Existing candidate metadata intentionally retains acceptance=false and release_qualified=false; those fields record scientific qualification, not whether the experimental package has been uploaded.
Known boundaries
- The input contract is bounded synthetic controller evidence. Do not relabel real executions to pass admission.
- No arbitrary-PR understanding, patch generation, autonomous engineering or production suitability is established.
- Missing observations remain unknown rather than fabricated zero counts.
- Internal hashes verify consistency, not source ownership, event absence or wall-clock ordering.
- Longer rollouts may contain unresolved/pruned probability mass; it is not silently renormalized away.
- Supported consumer horizons are 1, 2, 4 and 8; this release's installation smoke covers H1 only.
- The learned probabilities are not proven calibrated and do not authorize external actions.
- The real-patch exact comparator is a baseline, not learned prediction. Only one exposed development forecast was scored; no public API or Linux exact-execution qualification exists.
Benchmark and iteration
A future benchmark will compare identical pre-action evidence against native outcomes, with persistence/prior/source-rule baselines and Jev/Laya on shared questions. No Jev/Laya comparison has been run. Base versus fitted checkpoints will be separate arms. Public example replay is distinct from fresh prospective accuracy or independent transfer.
Later versions should pin source, schema, runtime, weights, data exposure and results together. This version remains available as a baseline, including its negative result. See CHANGELOG.md.
Author, attribution and license
Navneet Prabhakar β creator, maintainer and sole contributor to Laxmi.
Original components explicitly listed in RELEASE_SCOPE.md are released under Apache License 2.0. Preserve applicable attribution and license notices when redistributing; see NOTICE. Third-party dependencies retain their own terms. This release does not distribute the full corpus or private operational archives.
If you use Laxmi in research, please cite CITATION.cff. Citation is requested, not an additional license condition.