cortex / README.md
appvoid's picture
Clean model history
da49047
|
Raw History Blame
7.62 kB
metadata
library_name: transformers
pipeline_tag: text-generation
tags:
  - bet
  - byte-level
  - recurrent
  - looped

Cortex — SparkBET-9M

Target repository: appvoid/cortex. The current training model has 9,353,876 parameters: width 324, FFN 864, one prelude block, six physical recurrent body blocks, one coda block, 6 query heads / 2 KV heads, head dimension 54, rank-16 phase LoRA, two loop-level Hyper-Connection lanes, Deep-Delta body residuals, QK normalization, and continuous phase/stride conditioning. The byte vocabulary has 259 IDs (0–255 plus PAD/BOS/EOS), context is 1,024 IDs, and the maximum recurrent budget is L8.

Run

Import the notebook into Kaggle (two T4 GPUs) or a Modal notebook (A10/A10G or RTX PRO 6000), enable network access, provide HF_TOKEN, then Run All. Kaggle Secrets and environment variables are both supported. The same FP16 autocast + GradScaler precision policy is used on every supported CUDA family; larger-memory GPUs spend memory on larger physical microbatches before activation checkpointing is enabled.

The trainer runs continuously until interrupted. A fresh run warms up once and then keeps a nonzero constant learning rate. Interrupting the training cell requests a clean stop and waits for a completed optimizer update, local full-state save, Hugging Face full-state upload, metrics upload, and inference export. A hard kill can recover from the most recent durable checkpoint.

All project source remains embedded in readable notebook cells and is materialized by Run All. This includes data streams, the Cortex curriculum, model, trainer, full-state recovery, Hugging Face export, tokenizer/model wrappers, provenance files, and tests. Existing codec utility modules remain packaged for project continuity, but the current training registry contains no image or audio datasets.

Active training data

The active external language sources are:

  • appvoid/rewrite6
  • Ultra-FineWeb-L3 English multi-style
  • Ultra-FineWeb-L3 English QA
  • DCLM baseline 1.0
  • FineWeb-Edu score 2
  • FineMath 4+
  • appvoid/no-prompt-oasst
  • appvoid/no-prompt-openhermes
  • SmolLM-Corpus Cosmopedia v2
  • SmolLM-Corpus FineWeb-Edu-dedup

The procedural Cortex curriculum remains active as its own weighted source. Dataset revisions and shard lists are pinned in dataset_manifest.json; each source keeps deterministic cursor/epoch/rejection state in full checkpoints.

Objective and recurrent compute

Every successful optimizer update trains the exact L8 trajectory plus one exact auxiliary trajectory, keeping the same objective:

CE(L8) + 0.20 * CE(Lr), r in {1,...,7}.

The auxiliary depth follows a deterministic, repeating sustained progressive-data curriculum, not a rapid 28-update depth rotation. With aux_stage_base_updates=128 the L8+L1 stage receives 128 committed optimizer updates and batch draws, L8+L2 receives 256, then L3=384, L4=512, L5=640, L6=768, and L7=896; one curriculum spans 3,584 updates before repeating. Every update draws another global mixer batch (which can contain repeated finite-dataset samples after epochs, not guaranteed unique records). Hence L7 receives 7× the optimizer updates and example presentations allocated to L1, and L8 remains supervised in every update. Tune only aux_stage_base_updates to scale the entire schedule; the checkpoint persists this value and the curriculum origin. Depth is independent of microbatch count, DDP partition and overflow retries, and is reconstructed from the committed checkpoint step on resume. Each exact-budget trajectory receives its own continuous phase/stride coordinates; auxiliary losses are not snapshots from L8. Full BPTT is retained; activation checkpointing remains a GPU-memory fallback. More exposure encourages but does not guarantee monotonic validation performance.

The exact previous SparkBET notebook fingerprint and exact original pre-schedule fingerprint are explicitly approved for verified auxiliary-schedule migration. Migrated checkpoints begin their new data stages at L1 using their committed update as aux_curriculum_origin_step, while preserving model, optimizer, scaler, and data cursors. A subsequent resume restores the saved origin and stage budget instead of restarting L1. Other fingerprints are rejected rather than silently restarting. Set allow_verified_schedule_migration=False to forbid even these approved migrations. The new objective fingerprint is stored on the next checkpoint. BET2/LeWorldModel experimental code is included, but BET2_ABLATION_ENABLED=False by default; standard training does not create or train the sidecar.

Hardware policy

Run All identifies T4, A10/A10G, and RTX PRO 6000 explicitly. It first probes the largest divisible physical microbatch without activation checkpointing. If none fits with safety headroom, it retries with checkpointing. This keeps model/context/precision fixed while using larger-memory GPUs for larger physical batches. TF32 remains disabled so the numerical precision policy does not silently change when a run moves between supported GPU families.

Checkpoints, resume, and Hugging Face

Complete local checkpoint directories contain training.pt, metadata.json, and COMPLETE. They preserve model, AdamW, GradScaler, per-rank RNG, exact mixer/data cursors, update counters, and training contract. Local retention keeps three complete checkpoints.

Hugging Face full-state checkpoints are uploaded to a SparkBET-specific namespace inside appvoid/cortex so they cannot be confused with earlier architectures. Hub head retention keeps two complete SparkBET checkpoints; repository history is not rewritten. At startup the trainer compares verified local and Hub candidates newest-first and can recover across sessions. Upload failures preserve local state and are reported rather than deleting a completed checkpoint.

TensorBoard event files are uploaded independently to runs/; the model page exposes them under the repository TensorBoard view. A standalone Transformers export is uploaded periodically and on a clean stop. The export contains SafeTensors weights, exact configuration, byte tokenizer, model card, dataset manifest, and custom model code for trust_remote_code=True.

Hugging Face authentication

HF_TOKEN is read from Kaggle Secrets first on Kaggle and otherwise from the environment. It is never printed. The token needs read access to any private training source you use and write access to appvoid/cortex.

Inference

The Hub export supports:

from transformers import AutoTokenizer, AutoModelForCausalLM
import torch

repo = "appvoid/cortex"
tok = AutoTokenizer.from_pretrained(repo, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(repo, trust_remote_code=True).cuda().eval()
ids = torch.tensor([[257] + tok.encode("The next step is", add_special_tokens=False)], device="cuda")
with torch.inference_mode(), torch.autocast("cuda", dtype=torch.float16):
    result = model.generate(ids, max_new_tokens=200, do_sample=False, use_cache=False)
print(tok.decode(result[0], skip_special_tokens=True))

The current model does not implement a recurrent KV cache, so generation recomputes the retained context for each new byte.

Validation

The notebook runs source-integrity checks, architecture fingerprint/count checks, semantic/curriculum tests, data-stream tests, checkpoint recovery tests, Hugging Face export structure tests, and DDP equivalence checks before training. GPU capacity is validated with a real full-context L8+L7 backward/Adam preflight before the first committed update.