Papers
arxiv:2610.04002

Do Language Models Need a Trainable Input Embedding Table? Fixed Minimal Token Codes at 1.7B-Class Scale

Published on Oct 2
· Submitted by
Andrey
on Oct 7
Authors:

Abstract

A trainable input embedding table assigns each vocabulary item an independently adjustable vector. We investigate whether this token-specific parameterization is required for substantial language-modeling capability, or whether a shared Transformer can learn from fixed token identities. We compare three decoder-only language models trained from scratch with the same tokenizer, contextual backbone, untied output-head architecture, and training recipe, with a target budget of 100 billion prediction tokens per model. Their input interfaces are a learned table, canonical 16-bit token-ID codes, and one fixed invertible recoding over GF(2). The fixed codes are repeated to model width without an additional trainable input projection. Both fixed-code models acquire substantial capabilities: canonical codes achieve 52.40\% HellaSwag normalized accuracy, 70.51\% PIQA accuracy, and 42.75\% LAMBADA accuracy. The learned-input control performs better on several evaluations, including HellaSwag and LAMBADA, so these results establish viability rather than performance parity. The fixed interfaces remove 100.7 million trainable parameters, yielding 1.711B-parameter models, but parameter reduction is not the central result. These single-run experiments distinguish architectural necessity from empirical utility: independently trainable token-specific input vectors are not required for the observed capabilities. A fixed identity interface also provides a controlled setting for studying representation learning downstream of an immutable input, without establishing where particular capabilities are localized.

Community

Paper submitter

When we teach students about LLMs, explanations inherited from word2vec can leave the impression that tokens are inherently semantic units and that their input embeddings are where meaning resides. This work offers a concrete experiment for questioning that picture.

We replace the trainable input embedding table with fixed 16-dimensional binary vectors, deterministically repeated to model width, with no additional trainable input projection. This is a token-identity encoding, not weight quantization. At 1.7B-class scale, these models acquire substantial language-modeling capability, although the learned-input control performs better on several key benchmarks.

There is nothing linguistically special about sixteen dimensions: it is the minimum fixed-length binary width needed to distinguish the 49,152 tokens in our vocabulary. Wider binary vectors—for example, 32-dimensional ones—or suitably constructed, distinct sine/cosine-based token codes are plausible alternatives, but we do not test them here. Preserving token identity is a coding requirement; whether a network can learn effectively from a particular encoding is an empirical question.

The conceptual distinction is between identifying a token and learning to interpret and use token sequences in context. Tokenization defines the input units; it does not make each unit a self-contained piece of meaning.

Nor does removing the trainable input table establish that meaning simply moves into the first trainable layer. The experiment does not determine where semantic information is represented or how it is distributed across layers. Instead, fixing the input coordinates provides a controlled setting for investigating representation formation—without making the model automatically interpretable.

For students, the takeaway is that token identity, input representation, and meaning should not be conflated. Word-vector analogies can remain useful illustrations without becoming the lens through which all LLM computation is understood.

We report one training run per interface, and the output head remains trainable. Models and research artifacts: https://huggingface.co/collections/Bochkov/do-language-models-need-a-trainable-input-embedding-table

This is an automated message from the Librarian Bot. I found the following papers similar to this paper.

The following papers were recommended by the Semantic Scholar API

Please give a thumbs up to this comment if you found it helpful!

If you want recommendations for any Paper on Hugging Face checkout this Space

You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2610.04002
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 3

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2610.04002 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2610.04002 in a Space README.md to link it from this page.

Collections including this paper 1