Add Sentence Transformers usage

#1
by tomaarsen HF Staff - opened
Files changed (1) hide show
  1. README.md +34 -0
README.md CHANGED
@@ -2,6 +2,7 @@
2
  tags:
3
  - ColBERT
4
  - PyLate
 
5
  - sentence-transformers
6
  - sentence-similarity
7
  - embeddings
@@ -1034,6 +1035,39 @@ ColBERT(
1034
  ```
1035
 
1036
  ## Usage
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1037
  First install the PyLate library:
1038
 
1039
  ```bash
 
2
  tags:
3
  - ColBERT
4
  - PyLate
5
+ - multi-vector
6
  - sentence-transformers
7
  - sentence-similarity
8
  - embeddings
 
1035
  ```
1036
 
1037
  ## Usage
1038
+
1039
+ ### Sentence Transformers
1040
+
1041
+ This model can be used with [Sentence Transformers](https://www.sbert.net/) as a multi-vector (ColBERT-style late interaction) retriever via the `MultiVectorEncoder`:
1042
+
1043
+ ```bash
1044
+ pip install "sentence-transformers>=6.0.0"
1045
+ ```
1046
+
1047
+ ```python
1048
+ from sentence_transformers import MultiVectorEncoder
1049
+
1050
+ model = MultiVectorEncoder("lightonai/ColBERT-Zero-unsupervised-noprompts")
1051
+
1052
+ query = "What is the capital of France?"
1053
+ documents = [
1054
+ "Paris is the capital and largest city of France.",
1055
+ "Berlin is the capital of Germany.",
1056
+ ]
1057
+
1058
+ query_embeddings = model.encode_query(query)
1059
+ document_embeddings = model.encode_document(documents)
1060
+ print(query_embeddings.shape, document_embeddings[0].shape)
1061
+ # torch.Size([10, 128]) torch.Size([12, 128])
1062
+
1063
+ # MaxSim late-interaction scoring (higher is more relevant)
1064
+ scores = model.similarity(query_embeddings, document_embeddings)
1065
+ print(scores)
1066
+ # tensor([[7.2047, 5.4773]], device='cuda:0')
1067
+ ```
1068
+
1069
+ ### PyLate
1070
+
1071
  First install the PyLate library:
1072
 
1073
  ```bash