pplx-embed-v1-0.6b-acta-ita

*Developed at Infocube.*

Fine-tuned version of perplexity-ai/pplx-embed-v1-0.6b for the Italian administrative legal acts domain. The main goal is to obtain more reliable results in the mentioned domain compared to existing multilingual models.

Model Details

Developed by Infocube
Base model perplexity-ai/pplx-embed-v1-0.6b
Model Type Sentence Transformer
Base Model Max Sequence Length 32k tokens
Output Dimensionality 1024
Similarity Function Cosine Similarity
Language Italian

Usage

pip install -U sentence-transformers
from sentence_transformers import SentenceTransformer

model = SentenceTransformer("InfocubeSrl/pplx-embed-v1-0.6b-acta-it", trust_remote_code=True)

sentences = [
    "Nomina delegazione trattante di parte pubblica - CCNL Funzioni Locali 23/02/2026.",
    "Nomina delegazione trattante di parte pubblica - CCNL Funzioni locali 16 novembre 2022 – anno 2024.",
    "Costituzione della delegazione trattante di parte datoriale - CCNL Funzioni locali 16 novembre 2022. Anno 2025",
]
embeddings = model.encode(sentences)
similarities = model.similarity(embeddings, embeddings)
print(similarities)
# tensor([[1.0000, 0.8504, 0.6300],
#         [0.8504, 1.0000, 0.7438],
#         [0.6300, 0.7438, 1.0000]])

Evaluation

Held out on 1,000 triplets never seen in training, using TripletEvaluator (accuracy = fraction of triplets where the positive is closer to the anchor than the negative, cosine similarity).

Model Triplet accuracy cos(a, pos) cos(a, neg) Margin
jina-embeddings-v3 (baseline) 0.850 0.925 0.843 0.082
pplx-embed-v1-0.6b (base, untrained) 0.805 0.873 0.761 0.112
pplx-embed-v1-0.6b-acta-it 0.825 0.853 0.707 0.146
  • +2.0 accuracy points over the base model,
  • reaching 0.825 vs. 0.850 for jina-embeddings-v3 (2.5 points behind).
  • Negative similarity moves in the right direction: cos(anchor, negative) decreases
  • from 0.761 for the base model to 0.707 after fine-tuning, indicating
  • better separation from hard negatives.
  • On independent held-out topic groups, between-group separation improved by +32–46% across three difficulty levels; within-group cohesion also increased, so the model separates topics more aggressively without tightening them.

Training Dataset

  • Size: 9,522 training triplets, mined from ~134,000 Italian municipal act subject lines via k-means clustering (k=12,000, Euclidean) in the embedding space of jina-embeddings-v3: anchor = point closest to the cluster centroid, positive = its nearest neighbor inside the cluster, negative = the nearest point outside the cluster.

  • Columns: anchor, positive, negative

  • Samples:

    anchor positive negative
    ADOZIONE DELLA VARIANTE AL PROGRAMMA INTEGRATO DI INTERVENTO (PII) "VIMERCATE VECCHIO OSPEDALE - NORMA SPECIALE" E RELATIVO ATTO INTEGRATIVO DELLA CONVENZIONE URBANISTICA AI SENSI DELLA L.R. N. 12/2005 APPROVAZIONE DELLA VARIANTE AL PROGRAMMA INTEGRATO DI INTERVENTO (PII) "VIMERCATE VECCHIO OSPEDALE - NORMA SPECIALE" E RELATIVO ATTO INTEGRATIVO DELLA CONVENZIONE URBANISTICA AI SENSI DELLA L.R. N. 12/2005 E S.M.I. APPROVAZIONE PROTOCOLLO OPERATIVO PER LA GESTIONE DEI PROGETTI DI PUBBLICA UTILITA' (P.U.C.) SUL TERRITORIO DELL'AMBITO DI VIMERCATE
  • Loss: CachedMultipleNegativesRankingLoss (scale 20.0, cosine similarity, mini_batch_size 16)

Downloads last month
13
Safetensors
Model size
0.6B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for InfocubeSrl/pplx-embed-v1-0.6b-acta-ita

Finetuned
(7)
this model

Evaluation results