PathFinder Flan-T5 Large Second Try โ€” INT8 ONNX

INT8 ONNX encoder/decoder export stored with standard Optimum filenames for CPU deployment. The source directory is the PathFinderShip My Class/Second Try export.

Files

  • encoder_model.onnx
  • decoder_model.onnx
  • decoder_with_past_model.onnx
  • T5 tokenizer, model, and generation configuration files

The local source files used _int8 suffixes. Only the filenames were standardized for Optimum compatibility; the binary contents were not changed. The source-to-published mapping and SHA-256 values are recorded in ARTIFACT_SHA256.json.

Source-model evaluation

Metric Second Try LoRA
Chat token-F1 0.5216
RAG token-F1 0.8894
RAG exact match 0.7938

These scores belong to the source Second Try LoRA evaluation. They must not be interpreted as a separate ONNX parity benchmark. The complete comparison is included in evaluation/flan_retraining_results.json.

Retraining comparison

Usage

pip install "optimum[onnxruntime]" transformers
from optimum.onnxruntime import ORTModelForSeq2SeqLM
from transformers import AutoTokenizer

model_id = "Fatihaybasn/pathfinder-flan-t5-large-second-try-onnx-int8"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = ORTModelForSeq2SeqLM.from_pretrained(
    model_id,
    provider="CPUExecutionProvider",
)

Limitations

  • The export is intended for ONNX Runtime CPU inference.
  • Generation quality depends on the PathFinderShip prompt templates and decoding settings.
  • A clean-download smoke test is required before deleting the local source files.

Project documentation: PathFinderShip.

Downloads last month
19
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for Fatihaybasn/pathfinder-flan-t5-large-second-try-onnx-int8

Quantized
(23)
this model

Collection including Fatihaybasn/pathfinder-flan-t5-large-second-try-onnx-int8