MuXodious commited on
Commit
33bdeb0
·
verified ·
1 Parent(s): ab90b94
README.md CHANGED
@@ -1,199 +1,525 @@
1
  ---
2
  library_name: transformers
3
- tags: []
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
4
  ---
5
 
6
- # Model Card for Model ID
 
7
 
8
- <!-- Provide a quick summary of what the model is/does. -->
9
 
 
10
 
 
11
 
12
- ## Model Details
13
-
14
- ### Model Description
15
-
16
- <!-- Provide a longer summary of what this model is. -->
17
-
18
- This is the model card of a 🤗 transformers model that has been pushed on the Hub. This model card has been automatically generated.
19
-
20
- - **Developed by:** [More Information Needed]
21
- - **Funded by [optional]:** [More Information Needed]
22
- - **Shared by [optional]:** [More Information Needed]
23
- - **Model type:** [More Information Needed]
24
- - **Language(s) (NLP):** [More Information Needed]
25
- - **License:** [More Information Needed]
26
- - **Finetuned from model [optional]:** [More Information Needed]
27
-
28
- ### Model Sources [optional]
29
-
30
- <!-- Provide the basic links for the model. -->
31
-
32
- - **Repository:** [More Information Needed]
33
- - **Paper [optional]:** [More Information Needed]
34
- - **Demo [optional]:** [More Information Needed]
35
-
36
- ## Uses
37
-
38
- <!-- Address questions around how the model is intended to be used, including the foreseeable users of the model and those affected by the model. -->
39
-
40
- ### Direct Use
41
-
42
- <!-- This section is for the model use without fine-tuning or plugging into a larger ecosystem/app. -->
43
-
44
- [More Information Needed]
45
-
46
- ### Downstream Use [optional]
47
-
48
- <!-- This section is for the model use when fine-tuned for a task, or when plugged into a larger ecosystem/app -->
49
-
50
- [More Information Needed]
51
-
52
- ### Out-of-Scope Use
53
-
54
- <!-- This section addresses misuse, malicious use, and uses that the model will not work well for. -->
55
-
56
- [More Information Needed]
57
-
58
- ## Bias, Risks, and Limitations
59
-
60
- <!-- This section is meant to convey both technical and sociotechnical limitations. -->
61
-
62
- [More Information Needed]
63
-
64
- ### Recommendations
65
-
66
- <!-- This section is meant to convey recommendations with respect to the bias, risk, and technical limitations. -->
67
-
68
- Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. More information needed for further recommendations.
69
-
70
- ## How to Get Started with the Model
71
-
72
- Use the code below to get started with the model.
73
-
74
- [More Information Needed]
75
-
76
- ## Training Details
77
-
78
- ### Training Data
79
-
80
- <!-- This should link to a Dataset Card, perhaps with a short stub of information on what the training data is all about as well as documentation related to data pre-processing or additional filtering. -->
81
-
82
- [More Information Needed]
83
-
84
- ### Training Procedure
85
-
86
- <!-- This relates heavily to the Technical Specifications. Content here should link to that section when it is relevant to the training procedure. -->
87
-
88
- #### Preprocessing [optional]
89
-
90
- [More Information Needed]
91
-
92
-
93
- #### Training Hyperparameters
94
-
95
- - **Training regime:** [More Information Needed] <!--fp32, fp16 mixed precision, bf16 mixed precision, bf16 non-mixed precision, fp16 non-mixed precision, fp8 mixed precision -->
96
-
97
- #### Speeds, Sizes, Times [optional]
98
-
99
- <!-- This section provides information about throughput, start/end time, checkpoint size if relevant, etc. -->
100
-
101
- [More Information Needed]
102
-
103
- ## Evaluation
104
-
105
- <!-- This section describes the evaluation protocols and provides the results. -->
106
-
107
- ### Testing Data, Factors & Metrics
108
 
109
- #### Testing Data
 
 
 
 
110
 
111
- <!-- This should link to a Dataset Card if possible. -->
112
-
113
- [More Information Needed]
114
-
115
- #### Factors
116
-
117
- <!-- These are the things the evaluation is disaggregating by, e.g., subpopulations or domains. -->
118
-
119
- [More Information Needed]
120
-
121
- #### Metrics
122
-
123
- <!-- These are the evaluation metrics being used, ideally with a description of why. -->
124
-
125
- [More Information Needed]
126
-
127
- ### Results
128
-
129
- [More Information Needed]
130
-
131
- #### Summary
132
-
133
-
134
-
135
- ## Model Examination [optional]
136
-
137
- <!-- Relevant interpretability work for the model goes here -->
138
-
139
- [More Information Needed]
140
-
141
- ## Environmental Impact
142
-
143
- <!-- Total emissions (in grams of CO2eq) and additional considerations, such as electricity usage, go here. Edit the suggested text below accordingly -->
144
-
145
- Carbon emissions can be estimated using the [Machine Learning Impact calculator](https://mlco2.github.io/impact#compute) presented in [Lacoste et al. (2019)](https://arxiv.org/abs/1910.09700).
146
-
147
- - **Hardware Type:** [More Information Needed]
148
- - **Hours used:** [More Information Needed]
149
- - **Cloud Provider:** [More Information Needed]
150
- - **Compute Region:** [More Information Needed]
151
- - **Carbon Emitted:** [More Information Needed]
152
-
153
- ## Technical Specifications [optional]
154
-
155
- ### Model Architecture and Objective
156
-
157
- [More Information Needed]
158
-
159
- ### Compute Infrastructure
160
-
161
- [More Information Needed]
162
-
163
- #### Hardware
164
-
165
- [More Information Needed]
166
-
167
- #### Software
168
-
169
- [More Information Needed]
170
-
171
- ## Citation [optional]
172
-
173
- <!-- If there is a paper or blog post introducing the model, the APA and Bibtex information for that should go in this section. -->
174
-
175
- **BibTeX:**
176
-
177
- [More Information Needed]
178
-
179
- **APA:**
180
-
181
- [More Information Needed]
182
-
183
- ## Glossary [optional]
184
-
185
- <!-- If relevant, include terms and calculations in this section that can help readers understand the model or model card. -->
186
-
187
- [More Information Needed]
188
-
189
- ## More Information [optional]
190
-
191
- [More Information Needed]
192
-
193
- ## Model Card Authors [optional]
194
-
195
- [More Information Needed]
196
-
197
- ## Model Card Contact
198
-
199
- [More Information Needed]
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
  ---
2
  library_name: transformers
3
+ license: apache-2.0
4
+ pipeline_tag: image-text-to-text
5
+ language:
6
+ - en
7
+ - de
8
+ - fr
9
+ - es
10
+ - it
11
+ - pt
12
+ - hi
13
+ - ja
14
+ - ko
15
+ - zh
16
+ - ar
17
+ tags:
18
+ - vision
19
+ - multimodal
20
+ - conversational
21
+ - multilingual
22
+ - native-resolution
23
  ---
24
 
25
+ # North Micro Vision Instruct
26
+ ![North-Micro-Vision_Hero](https://cdn-uploads.huggingface.co/production/uploads/66d732effe6684fc16b12c28/qQQSd5Pldz30hvGMhNMq_.png)
27
 
28
+ North Micro Vision Instruct is a 2.4B-parameter open-weight vision-language model with native-resolution image support, released under the Apache 2.0 license. It is designed as a compact foundation for prototyping, task-specific fine-tuning, and specialized multimodal applications.
29
 
30
+ Developed by [Cohere](https://cohere.com/).
31
 
32
+ > **Technical deep dive:** Read the [North Micro Vision technical blog post](https://huggingface.co/blog/CohereLabs/meet-north-micro-vision-instruct) for architecture, training, and evaluation details.
33
 
34
+ ## Highlights
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
35
 
36
+ - Native-resolution image processing that preserves aspect ratios and fine visual detail.
37
+ - Broad image-understanding capabilities across VQA, captioning, grounding, OCR, charts, and documents.
38
+ - Multilingual and multi-image support.
39
+ - Compact 2.4B-parameter scale suited to customization and deployment experimentation.
40
+ - Apache 2.0-licensed model weights.
41
 
42
+ ## Model Details
43
+ | Property | Value |
44
+ | --- | --- |
45
+ | Model ID | `CohereLabs/North-Micro-Vision-Instruct` |
46
+ | Total parameters | 2.4B |
47
+ | Language model | 2B parameters |
48
+ | Vision encoder | 400M parameters; custom-trained starting from [SigLIP 2 SO400M](https://huggingface.co/google/siglip2-so400m-patch16-384) |
49
+ | Inputs | Interleaved text and images |
50
+ | Output | Text |
51
+ | Languages | English, German, French, Spanish, Italian, Portuguese, Hindi, Japanese, Korean, Chinese, Arabic, and more |
52
+ | Tokenizer vocabulary size | 262,144 |
53
+ | LM Backbone context window | 128K tokens |
54
+ | Multimodal training context | 8K tokens |
55
+ | Checkpoint precision | bfloat16 |
56
+ | License | Apache 2.0 |
57
+
58
+ The language backbone supports a 128K-token context window, but the validated operating range for multimodal prompts is up to 8K tokens. Longer multimodal contexts may rely on extrapolation and have not been benchmarked.
59
+
60
+ ## Quickstart
61
+
62
+ ### Installation
63
+
64
+ Install [PyTorch](https://pytorch.org/get-started/locally/) for your platform first. North Micro Vision requires Transformers 5.16.0, together with `accelerate` for automatic device placement and Pillow for image loading. Until Transformers 5.16.0 is released, install the runtime dependencies and Transformers from source:
65
+
66
+ ```bash
67
+ uv pip install accelerate pillow
68
+ uv pip install "git+https://github.com/huggingface/transformers.git"
69
+ ```
70
+
71
+ Once Transformers 5.16.0 is available on PyPI, install the released package with:
72
+
73
+ ```bash
74
+ uv pip install accelerate pillow "transformers==5.16.0"
75
+ ```
76
+
77
+ Flash Attention 2 is optional. On supported CUDA systems, install it with:
78
+
79
+ ```bash
80
+ uv pip install flash-attn --no-build-isolation
81
+ ```
82
+
83
+ If you do not use `uv`, replace `uv pip` with `pip` in the commands above.
84
+
85
+ ### Transformers
86
+
87
+ The following example loads an image from a URL and asks the model to describe it. Prompts can interleave text with one or more images; for text-only prompts, omit the image entries.
88
+
89
+ ```python
90
+ import torch
91
+ from transformers import AutoModelForImageTextToText, AutoProcessor
92
+
93
+ model_id = "CohereLabs/North-Micro-Vision-Instruct"
94
+
95
+ processor = AutoProcessor.from_pretrained(
96
+ model_id,
97
+ )
98
+ model = AutoModelForImageTextToText.from_pretrained(
99
+ model_id,
100
+ dtype="auto",
101
+ device_map="auto",
102
+ )
103
+
104
+ # To enable Flash Attention 2, load the model with the following settings:
105
+ # model = AutoModelForImageTextToText.from_pretrained(
106
+ # model_id,
107
+ # dtype=torch.bfloat16,
108
+ # attn_implementation="flash_attention_2",
109
+ # device_map="auto",
110
+ # )
111
+
112
+ image_url = "https://cdn-uploads.huggingface.co/production/uploads/66d732effe6684fc16b12c28/Io_5OCmftsmH-n158ZtPs.png"
113
+ messages = [
114
+ {
115
+ "role": "user",
116
+ "content": [
117
+ {"type": "image", "url": image_url},
118
+ {"type": "text", "text": "What do you see?"},
119
+ ],
120
+ }
121
+ ]
122
+
123
+ inputs = processor.apply_chat_template(
124
+ messages,
125
+ tokenize=True,
126
+ add_generation_prompt=True,
127
+ return_tensors="pt",
128
+ return_dict=True,
129
+ ).to(model.device)
130
+
131
+ outputs = model.generate(
132
+ **inputs,
133
+ max_new_tokens=128,
134
+ do_sample=True,
135
+ temperature=0.7,
136
+ top_p=0.8,
137
+ top_k=20,
138
+ )
139
+
140
+ generated_ids = [
141
+ output_ids[len(input_ids) :]
142
+ for input_ids, output_ids in zip(inputs.input_ids, outputs)
143
+ ]
144
+ response = processor.batch_decode(
145
+ generated_ids,
146
+ skip_special_tokens=True,
147
+ clean_up_tokenization_spaces=False,
148
+ )[0]
149
+ print(response)
150
+ ```
151
+
152
+ The example uses the recommended Transformers sampling settings. For deterministic output, set `do_sample=False` and omit `temperature`, `top_p`, and `top_k`.
153
+
154
+ ### Grounding Coordinates
155
+
156
+ Bounding boxes are returned as `[x1, y1, x2, y2]` on a normalized 0–1000 scale. Map them back to the original image by scaling each axis:
157
+
158
+ ```python
159
+ x1_px = x1 / 1000 * image_width
160
+ y1_px = y1 / 1000 * image_height
161
+ x2_px = x2 / 1000 * image_width
162
+ y2_px = y2 / 1000 * image_height
163
+ ```
164
+
165
+ ### vLLM
166
+
167
+ Public vLLM support is coming soon. Until it is available, use Transformers as shown above. The recommended vLLM settings will be:
168
+
169
+ ```python
170
+ temperature = 0.7
171
+ top_p = 0.8
172
+ top_k = 20
173
+ min_p = 0.0
174
+ presence_penalty = 1.5
175
+ repetition_penalty = 1.0
176
+ ```
177
+
178
+ ## Intended Use
179
+
180
+ North Micro Vision Instruct is intended for research and development use cases such as:
181
+
182
+ - Prototyping and task-specific fine-tuning.
183
+ - General visual question answering and image captioning.
184
+ - Multilingual and multi-image understanding.
185
+ - Visual grounding and spatial understanding.
186
+ - OCR, chart and document understanding, and structured information extraction.
187
+
188
+ ## Limitations
189
+
190
+ - The model is intended as a compact foundation for customization rather than a replacement for larger general-purpose chat assistants.
191
+ - It is not a reasoning model and has limited math and code-generation capabilities.
192
+ - Tool calling and agentic workflows are not supported.
193
+ - System prompts are not recommended because the model was not trained with them, although the chat template accepts the `system` role.
194
+ - Multimodal training used an 8K-token context; longer contexts have not been validated.
195
+ - Native-resolution inputs can increase memory use and latency as image dimensions grow.
196
+
197
+ ## Benchmark Results
198
+
199
+ The complete comparison is provided below. We ran vision-language and text-only evaluations with [VLMEvalKit](https://github.com/open-compass/vlmevalkit), capping generation at 1,024 tokens; see the [technical blog post](https://huggingface.co/blog/CohereLabs/meet-north-micro-vision-instruct) for the full methodology.
200
+
201
+ <table>
202
+ <thead>
203
+ <tr>
204
+ <th></th>
205
+ <th style="font-weight: bold; background-color: rgba(127, 127, 127, 0.08);">North-Micro-Vision-Instruct</th>
206
+ <th>Ministral-3-3B-Instruct</th>
207
+ <th>LFM2.5-VL-1.6B</th>
208
+ <th>Phi-3.5-vision-instruct</th>
209
+ <th>Gemma-4-E2B-it</th>
210
+ <th>Qwen3-VL-2B-Instruct</th>
211
+ <th>Qwen3.5-2B-Instruct</th>
212
+ <th>SmolVLM2.2B</th>
213
+ </tr>
214
+ </thead>
215
+ <tbody>
216
+ <tr>
217
+ <th style="text-align: left; font-weight: normal;">Size</th>
218
+ <td style="background-color: rgba(127, 127, 127, 0.08);">2.4B</td>
219
+ <td>3.8B</td>
220
+ <td>1.6B</td>
221
+ <td>4.2B</td>
222
+ <td>5.1B</td>
223
+ <td>2.2B</td>
224
+ <td>2.1B</td>
225
+ <td>2.2B</td>
226
+ </tr>
227
+ <tr>
228
+ <th style="text-align: left; font-weight: normal;">License</th>
229
+ <td style="background-color: rgba(127, 127, 127, 0.08);">Apache 2.0</td>
230
+ <td>Apache 2.0</td>
231
+ <td>LFM v1.0</td>
232
+ <td>MIT</td>
233
+ <td>Apache 2.0</td>
234
+ <td>Apache 2.0</td>
235
+ <td>Apache 2.0</td>
236
+ <td>Apache 2.0</td>
237
+ </tr>
238
+ <tr>
239
+ <th colspan="9" style="text-align: left; background-color: rgba(127, 127, 127, 0.16);">General VQA</th>
240
+ </tr>
241
+ <tr>
242
+ <th style="text-align: left; font-weight: normal;">MMBench<sub>DEV_EN_V11</sub></th>
243
+ <td style="background-color: rgba(127, 127, 127, 0.08);">0.687</td>
244
+ <td>0.692</td>
245
+ <td>0.696</td>
246
+ <td>0.731</td>
247
+ <td>0.693</td>
248
+ <td>0.744</td>
249
+ <td>0.760</td>
250
+ <td>0.674</td>
251
+ </tr>
252
+ <tr>
253
+ <th style="text-align: left; font-weight: normal;">MMStar</th>
254
+ <td style="background-color: rgba(127, 127, 127, 0.08);">0.518</td>
255
+ <td>0.531</td>
256
+ <td>0.508</td>
257
+ <td>0.495</td>
258
+ <td>0.529</td>
259
+ <td>0.506</td>
260
+ <td>0.614</td>
261
+ <td>0.460</td>
262
+ </tr>
263
+ <tr>
264
+ <th style="text-align: left; font-weight: normal;">RealWorldQA</th>
265
+ <td style="background-color: rgba(127, 127, 127, 0.08);">0.622</td>
266
+ <td>0.583</td>
267
+ <td>0.642</td>
268
+ <td>0.580</td>
269
+ <td>0.507</td>
270
+ <td>0.646</td>
271
+ <td>0.693</td>
272
+ <td>0.567</td>
273
+ </tr>
274
+ <tr>
275
+ <th style="text-align: left; font-weight: normal;">GQA<sub>TestDev_Balanced</sub></th>
276
+ <td style="background-color: rgba(127, 127, 127, 0.08);">0.574</td>
277
+ <td>0.544</td>
278
+ <td>0.395</td>
279
+ <td>0.650</td>
280
+ <td>0.387</td>
281
+ <td>0.572</td>
282
+ <td>0.539</td>
283
+ <td>0.000<sup>&Dagger;</sup></td>
284
+ </tr>
285
+ <tr>
286
+ <th colspan="9" style="text-align: left; background-color: rgba(127, 127, 127, 0.16);">Multilingual</th>
287
+ </tr>
288
+ <tr>
289
+ <th style="text-align: left; font-weight: normal;">MTL<sub>MMBench_DEV</sub></th>
290
+ <td style="background-color: rgba(127, 127, 127, 0.08);">0.636</td>
291
+ <td>0.674</td>
292
+ <td>0.623</td>
293
+ <td>0.619</td>
294
+ <td>0.648</td>
295
+ <td>0.664</td>
296
+ <td>0.669</td>
297
+ <td>0.454</td>
298
+ </tr>
299
+ <tr>
300
+ <th style="text-align: left; font-weight: normal;">MMMB</th>
301
+ <td style="background-color: rgba(127, 127, 127, 0.08);">0.728</td>
302
+ <td>0.734</td>
303
+ <td>0.717</td>
304
+ <td>0.686</td>
305
+ <td>0.743</td>
306
+ <td>0.723</td>
307
+ <td>0.745</td>
308
+ <td>0.577</td>
309
+ </tr>
310
+ <tr>
311
+ <th colspan="9" style="text-align: left; background-color: rgba(127, 127, 127, 0.16);">Multi-image</th>
312
+ </tr>
313
+ <tr>
314
+ <th style="text-align: left; font-weight: normal;">BLINK</th>
315
+ <td style="background-color: rgba(127, 127, 127, 0.08);">0.527</td>
316
+ <td>0.471</td>
317
+ <td>0.484</td>
318
+ <td>0.561</td>
319
+ <td>0.468</td>
320
+ <td>0.514</td>
321
+ <td>0.563</td>
322
+ <td>0.420</td>
323
+ </tr>
324
+ <tr>
325
+ <th colspan="9" style="text-align: left; background-color: rgba(127, 127, 127, 0.16);">Chart / Document / OCR</th>
326
+ </tr>
327
+ <tr>
328
+ <th style="text-align: left; font-weight: normal;">ChartQA<sub>Test</sub></th>
329
+ <td style="background-color: rgba(127, 127, 127, 0.08);">0.808</td>
330
+ <td>0.791</td>
331
+ <td>0.739</td>
332
+ <td>0.821</td>
333
+ <td>0.422</td>
334
+ <td>0.693</td>
335
+ <td>0.775</td>
336
+ <td>0.682</td>
337
+ </tr>
338
+ <tr>
339
+ <th style="text-align: left; font-weight: normal;">DocVQA<sub>VAL</sub></th>
340
+ <td style="background-color: rgba(127, 127, 127, 0.08);">0.921</td>
341
+ <td>0.896</td>
342
+ <td>0.877</td>
343
+ <td>0.860</td>
344
+ <td>0.732</td>
345
+ <td>0.825</td>
346
+ <td>0.926</td>
347
+ <td>0.799</td>
348
+ </tr>
349
+ <tr>
350
+ <th style="text-align: left; font-weight: normal;">InfoVQA<sub>VAL</sub></th>
351
+ <td style="background-color: rgba(127, 127, 127, 0.08);">0.652</td>
352
+ <td>0.589</td>
353
+ <td>0.627</td>
354
+ <td>0.561</td>
355
+ <td>0.380</td>
356
+ <td>0.622</td>
357
+ <td>0.731</td>
358
+ <td>0.383</td>
359
+ </tr>
360
+ <tr>
361
+ <th style="text-align: left; font-weight: normal;">OCRBench<sub>v2_en</sub></th>
362
+ <td style="background-color: rgba(127, 127, 127, 0.08);">0.367</td>
363
+ <td>0.414</td>
364
+ <td>0.415</td>
365
+ <td>0.339</td>
366
+ <td>0.435</td>
367
+ <td>0.417</td>
368
+ <td>0.481</td>
369
+ <td>0.304</td>
370
+ </tr>
371
+ <tr>
372
+ <th style="text-align: left; font-weight: normal;">OCRBench</th>
373
+ <td style="background-color: rgba(127, 127, 127, 0.08);">0.792</td>
374
+ <td>0.735</td>
375
+ <td>0.802</td>
376
+ <td>0.642</td>
377
+ <td>0.719</td>
378
+ <td>0.751</td>
379
+ <td>0.861</td>
380
+ <td>0.727</td>
381
+ </tr>
382
+ <tr>
383
+ <th style="text-align: left; font-weight: normal;">AI2D_TEST</th>
384
+ <td style="background-color: rgba(127, 127, 127, 0.08);">0.775</td>
385
+ <td>0.741</td>
386
+ <td>0.728</td>
387
+ <td>0.790</td>
388
+ <td>0.712</td>
389
+ <td>0.713</td>
390
+ <td>0.752</td>
391
+ <td>0.697</td>
392
+ </tr>
393
+ <tr>
394
+ <th style="text-align: left; font-weight: normal;">CharXiv<sub>DQ</sub></th>
395
+ <td style="background-color: rgba(127, 127, 127, 0.08);">0.600</td>
396
+ <td>0.766</td>
397
+ <td>0.516</td>
398
+ <td>0.637</td>
399
+ <td>0.751</td>
400
+ <td>0.595</td>
401
+ <td>0.761</td>
402
+ <td>0.482</td>
403
+ </tr>
404
+ <tr>
405
+ <th colspan="9" style="text-align: left; background-color: rgba(127, 127, 127, 0.16);">STEM</th>
406
+ </tr>
407
+ <tr>
408
+ <th style="text-align: left; font-weight: normal;">MMMU<sub>DEV_VAL</sub></th>
409
+ <td style="background-color: rgba(127, 127, 127, 0.08);">0.329</td>
410
+ <td>0.508</td>
411
+ <td>0.380</td>
412
+ <td>0.432</td>
413
+ <td>0.477</td>
414
+ <td>0.379</td>
415
+ <td>0.474</td>
416
+ <td>0.399</td>
417
+ </tr>
418
+ <tr>
419
+ <th colspan="9" style="text-align: left; background-color: rgba(127, 127, 127, 0.16);">Grounding / Counting</th>
420
+ </tr>
421
+ <tr>
422
+ <th style="text-align: left; font-weight: normal;">RefCOCO<sub>avg</sub><sup>&dagger;</sup></th>
423
+ <td style="background-color: rgba(127, 127, 127, 0.08);">0.732</td>
424
+ <td>0.317</td>
425
+ <td>0.581</td>
426
+ <td>0.451</td>
427
+ <td>0.084</td>
428
+ <td>0.304</td>
429
+ <td>0.785</td>
430
+ <td>0.018</td>
431
+ </tr>
432
+ <tr>
433
+ <th style="text-align: left; font-weight: normal;">CountBench</th>
434
+ <td style="background-color: rgba(127, 127, 127, 0.08);">0.725</td>
435
+ <td>0.737</td>
436
+ <td>0.910</td>
437
+ <td>0.645</td>
438
+ <td>0.534</td>
439
+ <td>0.848</td>
440
+ <td>0.805</td>
441
+ <td>0.764</td>
442
+ </tr>
443
+ <tr>
444
+ <th colspan="9" style="text-align: left; background-color: rgba(127, 127, 127, 0.16);">Robustness / Hallucination</th>
445
+ </tr>
446
+ <tr>
447
+ <th style="text-align: left; font-weight: normal;">HallusionBench</th>
448
+ <td style="background-color: rgba(127, 127, 127, 0.08);">0.615</td>
449
+ <td>0.652</td>
450
+ <td>0.601</td>
451
+ <td>0.585</td>
452
+ <td>0.598</td>
453
+ <td>0.673</td>
454
+ <td>0.655</td>
455
+ <td>0.600</td>
456
+ </tr>
457
+ <tr>
458
+ <th colspan="9" style="text-align: left; background-color: rgba(127, 127, 127, 0.16);">Text</th>
459
+ </tr>
460
+ <tr>
461
+ <th style="text-align: left; font-weight: normal;">MMLU<sub>test</sub></th>
462
+ <td style="background-color: rgba(127, 127, 127, 0.08);">0.504</td>
463
+ <td>0.660</td>
464
+ <td>0.464</td>
465
+ <td>0.355</td>
466
+ <td>0.692</td>
467
+ <td>0.630</td>
468
+ <td>0.543</td>
469
+ <td>0.084</td>
470
+ </tr>
471
+ <tr>
472
+ <th style="text-align: left; font-weight: normal;">MMLU-Pro<sub>test</sub></th>
473
+ <td style="background-color: rgba(127, 127, 127, 0.08);">0.307</td>
474
+ <td>0.475</td>
475
+ <td>0.199</td>
476
+ <td>0.286</td>
477
+ <td>0.441</td>
478
+ <td>0.428</td>
479
+ <td>0.298</td>
480
+ <td>0.099</td>
481
+ </tr>
482
+ <tr>
483
+ <th style="text-align: left; font-weight: normal;">Multi-If</th>
484
+ <td style="background-color: rgba(127, 127, 127, 0.08);">0.373</td>
485
+ <td>0.470</td>
486
+ <td>0.443</td>
487
+ <td>0.304</td>
488
+ <td>0.687</td>
489
+ <td>0.523</td>
490
+ <td>0.464</td>
491
+ <td>0.236</td>
492
+ </tr>
493
+ <tr>
494
+ <th style="text-align: left; font-weight: normal;">IFEval</th>
495
+ <td style="background-color: rgba(127, 127, 127, 0.08);">0.749</td>
496
+ <td>0.725</td>
497
+ <td>0.776</td>
498
+ <td>0.543</td>
499
+ <td>0.869</td>
500
+ <td>0.734</td>
501
+ <td>0.679</td>
502
+ <td>0.501</td>
503
+ </tr>
504
+ </tbody>
505
+ </table>
506
+ <p><small><sup>&dagger;</sup> Averaged over RefCOCO_val, RefCOCO_testA, RefCOCO_testB, RefCOCO+_val, RefCOCO+_testA, RefCOCO+_testB, RefCOCOg_val, RefCOCOg_test.</small></p>
507
+ <p><small><sup>&Dagger;</sup> SmolVLM2.2B's GQA output was scored as 0.000 under VLMEvalKit's answer-extraction rules.</small></p>
508
+
509
+
510
+
511
+ ## Citation
512
+
513
+ ```bibtex
514
+ @misc{cohere_north_micro_vision_instruct,
515
+ title = {{North Micro Vision}: A 2.4B Native-Resolution Vision-Language Model},
516
+ url = {https://huggingface.co/blog/CohereLabs/meet-north-micro-vision-instruct},
517
+ author = {{Team Cohere}},
518
+ month = {August},
519
+ year = {2026}
520
+ }
521
+ ```
522
+
523
+ ## Contact
524
+
525
+ For errors or questions about this model card, contact [Cohere Labs](mailto:labs@cohere.com).
config.json CHANGED
@@ -1,27 +1,72 @@
1
  {
 
2
  "architectures": [
3
  "CohereCompassForConditionalGeneration"
4
  ],
 
 
5
  "dtype": "bfloat16",
6
- "eos_token_id": 255001,
7
- "fusion_config": {
8
- "patch_embeddings": true
 
 
9
  },
10
- "image_token_id": 255031,
11
- "model_type": "cohere_compass",
12
- "pad_token_id": 0,
 
 
13
  "text_config": {
14
- "attention_bias": false,
15
- "attention_dropout": 0.0,
16
- "bos_token_id": 2,
17
  "dtype": "bfloat16",
18
- "eos_token_id": 255001,
19
- "head_dim": 128,
20
- "hidden_act": "silu",
 
 
 
 
 
 
 
 
 
21
  "hidden_size": 2048,
22
- "initializer_range": 0.02,
23
  "intermediate_size": 6144,
 
 
 
 
 
 
 
24
  "layer_norm_eps": 1e-05,
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
25
  "layer_types": [
26
  "sliding_attention",
27
  "sliding_attention",
@@ -52,60 +97,62 @@
52
  "sliding_attention",
53
  "full_attention"
54
  ],
55
- "logit_scale": 0.25,
56
- "max_position_embeddings": 500000,
57
- "model_type": "cohere_compass_text",
58
- "num_attention_heads": 16,
59
- "num_hidden_layers": 28,
60
- "num_key_value_heads": 8,
61
- "pad_token_id": 0,
62
- "pooling": null,
63
- "rope_parameters": {
64
- "full_attention": null,
65
- "rope_theta": 10000.0,
66
- "rope_type": "default",
67
- "sliding_attention": {
68
- "mrope_interleaved": true,
69
- "mrope_section": [
70
- 24,
71
- 20,
72
- 20
73
- ],
74
- "rope_theta": 50000,
75
- "rope_type": "default"
76
- }
77
- },
78
  "score_shift_a": null,
79
  "score_shift_b": null,
80
- "sliding_window": 4096,
81
- "tie_word_embeddings": true,
82
- "use_cache": true,
83
- "vocab_size": 262144
 
84
  },
85
- "tie_word_embeddings": true,
86
- "transformers_version": "5.16.0.dev0",
87
- "video_token_id": 255032,
88
  "vision_config": {
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
89
  "deepstack_visual_indexes": [
90
  8,
91
  16,
92
  24
93
  ],
94
- "depth": 27,
95
- "dtype": "bfloat16",
96
- "hidden_act": "gelu_pytorch_tanh",
97
- "hidden_size": 1152,
98
- "in_channels": 3,
99
  "initializer_range": 0.02,
100
- "intermediate_size": 4304,
101
  "model_type": "cohere_compass_vision",
102
- "num_heads": 16,
103
- "num_position_embeddings": 2304,
104
- "out_hidden_size": 2048,
105
- "patch_size": 16,
106
- "spatial_merge_size": 2,
107
- "temporal_patch_size": 2
108
  },
 
 
 
109
  "vision_end_token_id": 255029,
110
- "vision_start_token_id": 255028
111
- }
 
 
 
 
 
 
1
  {
2
+ "transformers_version": "5.15.0.dev0",
3
  "architectures": [
4
  "CohereCompassForConditionalGeneration"
5
  ],
6
+ "output_hidden_states": false,
7
+ "return_dict": true,
8
  "dtype": "bfloat16",
9
+ "chunk_size_feed_forward": 0,
10
+ "is_encoder_decoder": false,
11
+ "id2label": {
12
+ "0": "LABEL_0",
13
+ "1": "LABEL_1"
14
  },
15
+ "label2id": {
16
+ "LABEL_0": 0,
17
+ "LABEL_1": 1
18
+ },
19
+ "problem_type": null,
20
  "text_config": {
21
+ "architectures": null,
22
+ "output_hidden_states": false,
23
+ "return_dict": true,
24
  "dtype": "bfloat16",
25
+ "chunk_size_feed_forward": 0,
26
+ "is_encoder_decoder": false,
27
+ "id2label": {
28
+ "0": "LABEL_0",
29
+ "1": "LABEL_1"
30
+ },
31
+ "label2id": {
32
+ "LABEL_0": 0,
33
+ "LABEL_1": 1
34
+ },
35
+ "problem_type": null,
36
+ "vocab_size": 262144,
37
  "hidden_size": 2048,
 
38
  "intermediate_size": 6144,
39
+ "logit_scale": 0.25,
40
+ "num_hidden_layers": 28,
41
+ "num_attention_heads": 16,
42
+ "num_key_value_heads": 8,
43
+ "hidden_act": "silu",
44
+ "max_position_embeddings": 500000,
45
+ "initializer_range": 0.02,
46
  "layer_norm_eps": 1e-05,
47
+ "use_cache": true,
48
+ "pad_token_id": 0,
49
+ "bos_token_id": 2,
50
+ "eos_token_id": 255001,
51
+ "tie_word_embeddings": true,
52
+ "rope_parameters": {
53
+ "sliding_attention": {
54
+ "mrope_interleaved": true,
55
+ "mrope_section": [
56
+ 24,
57
+ 20,
58
+ 20
59
+ ],
60
+ "rope_type": "default",
61
+ "rope_theta": 50000
62
+ },
63
+ "full_attention": null,
64
+ "rope_theta": 10000.0,
65
+ "rope_type": "default"
66
+ },
67
+ "attention_bias": false,
68
+ "attention_dropout": 0.0,
69
+ "sliding_window": 4096,
70
  "layer_types": [
71
  "sliding_attention",
72
  "sliding_attention",
 
97
  "sliding_attention",
98
  "full_attention"
99
  ],
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
100
  "score_shift_a": null,
101
  "score_shift_b": null,
102
+ "pooling": null,
103
+ "head_dim": 128,
104
+ "_name_or_path": "",
105
+ "model_type": "cohere_compass_text",
106
+ "output_attentions": false
107
  },
 
 
 
108
  "vision_config": {
109
+ "architectures": null,
110
+ "output_hidden_states": false,
111
+ "return_dict": true,
112
+ "dtype": "bfloat16",
113
+ "chunk_size_feed_forward": 0,
114
+ "is_encoder_decoder": false,
115
+ "id2label": {
116
+ "0": "LABEL_0",
117
+ "1": "LABEL_1"
118
+ },
119
+ "label2id": {
120
+ "LABEL_0": 0,
121
+ "LABEL_1": 1
122
+ },
123
+ "problem_type": null,
124
+ "depth": 27,
125
+ "hidden_size": 1152,
126
+ "hidden_act": "gelu_pytorch_tanh",
127
+ "intermediate_size": 4304,
128
+ "num_heads": 16,
129
+ "in_channels": 3,
130
+ "patch_size": 16,
131
+ "spatial_merge_size": 2,
132
+ "temporal_patch_size": 2,
133
+ "out_hidden_size": 2048,
134
+ "num_position_embeddings": 2304,
135
  "deepstack_visual_indexes": [
136
  8,
137
  16,
138
  24
139
  ],
 
 
 
 
 
140
  "initializer_range": 0.02,
141
+ "_name_or_path": "",
142
  "model_type": "cohere_compass_vision",
143
+ "output_attentions": false
144
+ },
145
+ "fusion_config": {
146
+ "patch_embeddings": true
 
 
147
  },
148
+ "image_token_id": 255031,
149
+ "video_token_id": 255032,
150
+ "vision_start_token_id": 255028,
151
  "vision_end_token_id": 255029,
152
+ "tie_word_embeddings": true,
153
+ "_name_or_path": "",
154
+ "eos_token_id": 255001,
155
+ "pad_token_id": 0,
156
+ "model_type": "cohere_compass",
157
+ "output_attentions": false
158
+ }
generation_config.json CHANGED
@@ -1,12 +1,12 @@
1
  {
2
  "bos_token_id": 2,
3
- "do_sample": true,
4
  "eos_token_id": [
5
  255001
6
  ],
7
  "pad_token_id": 0,
 
8
  "temperature": 0.7,
9
- "top_k": 20,
10
  "top_p": 0.8,
11
- "transformers_version": "5.16.0.dev0"
12
  }
 
 
1
  {
2
  "bos_token_id": 2,
 
3
  "eos_token_id": [
4
  255001
5
  ],
6
  "pad_token_id": 0,
7
+ "do_sample": true,
8
  "temperature": 0.7,
 
9
  "top_p": 0.8,
10
+ "top_k": 20
11
  }
12
+
preprocessor_config.json ADDED
@@ -0,0 +1,37 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "crop_size": null,
3
+ "data_format": "channels_first",
4
+ "default_to_square": false,
5
+ "device": null,
6
+ "disable_grouping": null,
7
+ "do_center_crop": null,
8
+ "do_convert_rgb": true,
9
+ "do_normalize": true,
10
+ "do_pad": null,
11
+ "do_rescale": true,
12
+ "do_resize": true,
13
+ "image_mean": [
14
+ 0.5,
15
+ 0.5,
16
+ 0.5
17
+ ],
18
+ "image_std": [
19
+ 0.5,
20
+ 0.5,
21
+ 0.5
22
+ ],
23
+ "input_data_format": null,
24
+ "merge_size": 2,
25
+ "pad_size": null,
26
+ "patch_size": 16,
27
+ "processor_class": "CohereCompassProcessor",
28
+ "resample": 3,
29
+ "rescale_factor": 0.00392156862745098,
30
+ "return_tensors": null,
31
+ "size": {
32
+ "longest_edge": 16777216,
33
+ "shortest_edge": 65536
34
+ },
35
+ "temporal_patch_size": 2,
36
+ "image_processor_type": "CohereCompassImageProcessor"
37
+ }
special_tokens_map.json ADDED
@@ -0,0 +1,37 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "additional_special_tokens": [
3
+ "<|VISION_START|>",
4
+ "<|IMAGE_PAD|>",
5
+ "<|VISION_END|>",
6
+ "<|VISION_PAD|>",
7
+ "<|VIDEO_PAD|>"
8
+ ],
9
+ "bos_token": {
10
+ "content": "<BOS_TOKEN>",
11
+ "lstrip": false,
12
+ "normalized": false,
13
+ "rstrip": false,
14
+ "single_word": false
15
+ },
16
+ "eos_token": {
17
+ "content": "<|END_OF_TURN_TOKEN|>",
18
+ "lstrip": false,
19
+ "normalized": false,
20
+ "rstrip": false,
21
+ "single_word": false
22
+ },
23
+ "pad_token": {
24
+ "content": "<PAD>",
25
+ "lstrip": false,
26
+ "normalized": false,
27
+ "rstrip": false,
28
+ "single_word": false
29
+ },
30
+ "unk_token": {
31
+ "content": "<UNK>",
32
+ "lstrip": false,
33
+ "normalized": false,
34
+ "rstrip": false,
35
+ "single_word": false
36
+ }
37
+ }
tokenizer.json CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:4594953eca27aa1122252e8a61a91a5b4941c92aba78d3811c8fe9b1b95d7f68
3
- size 19550475
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:6fcc5292908e0c8ad1400c67fe9413825c656486fe4190dce7d8d62c38bfbbf6
3
+ size 19550662
tokenizer_config.json CHANGED
@@ -1,29 +1,333 @@
1
  {
 
 
2
  "add_prefix_space": false,
3
- "backend": "tokenizers",
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
4
  "bos_token": "<BOS_TOKEN>",
5
  "clean_up_tokenization_spaces": false,
6
- "cls_token": "<CLS>",
7
  "eos_token": "<|END_OF_TURN_TOKEN|>",
8
- "errors": "replace",
9
  "image_token": "<|IMAGE_PAD|>",
10
- "is_local": true,
11
  "legacy": true,
12
- "local_files_only": false,
13
- "mask_token": "<MASK_TOKEN>",
14
  "max_pixels": 3868706,
 
15
  "min_pixels": 16384,
16
  "model_max_length": 1000000000000000019884624838656,
17
- "model_specific_special_tokens": {
18
- "image_token": "<|IMAGE_PAD|>",
19
- "video_token": "<|VIDEO_PAD|>",
20
- "vision_end_token": "<|VISION_END|>",
21
- "vision_start_token": "<|VISION_START|>"
22
- },
23
  "pad_token": "<PAD>",
24
  "padding_side": "right",
25
  "processor_class": "CohereCompassProcessor",
26
- "sep_token": "<SEP>",
27
  "sp_model_kwargs": {},
28
  "spaces_between_special_tokens": false,
29
  "tokenizer_class": "CohereTokenizer",
@@ -31,5 +335,6 @@
31
  "use_default_system_prompt": false,
32
  "video_token": "<|VIDEO_PAD|>",
33
  "vision_end_token": "<|VISION_END|>",
34
- "vision_start_token": "<|VISION_START|>"
 
35
  }
 
1
  {
2
+ "add_bos_token": false,
3
+ "add_eos_token": false,
4
  "add_prefix_space": false,
5
+ "added_tokens_decoder": {
6
+ "0": {
7
+ "content": "<PAD>",
8
+ "lstrip": false,
9
+ "normalized": false,
10
+ "rstrip": false,
11
+ "single_word": false,
12
+ "special": true
13
+ },
14
+ "1": {
15
+ "content": "<MASK_TOKEN>",
16
+ "lstrip": false,
17
+ "normalized": false,
18
+ "rstrip": false,
19
+ "single_word": false,
20
+ "special": true
21
+ },
22
+ "2": {
23
+ "content": "<BOS_TOKEN>",
24
+ "lstrip": false,
25
+ "normalized": false,
26
+ "rstrip": false,
27
+ "single_word": false,
28
+ "special": true
29
+ },
30
+ "3": {
31
+ "content": "<EOS_TOKEN>",
32
+ "lstrip": false,
33
+ "normalized": false,
34
+ "rstrip": false,
35
+ "single_word": false,
36
+ "special": true
37
+ },
38
+ "4": {
39
+ "content": "<UNK>",
40
+ "lstrip": false,
41
+ "normalized": false,
42
+ "rstrip": false,
43
+ "single_word": false,
44
+ "special": true
45
+ },
46
+ "255000": {
47
+ "content": "<|START_OF_TURN_TOKEN|>",
48
+ "lstrip": false,
49
+ "normalized": false,
50
+ "rstrip": false,
51
+ "single_word": false,
52
+ "special": false
53
+ },
54
+ "255001": {
55
+ "content": "<|END_OF_TURN_TOKEN|>",
56
+ "lstrip": false,
57
+ "normalized": false,
58
+ "rstrip": false,
59
+ "single_word": false,
60
+ "special": true
61
+ },
62
+ "255002": {
63
+ "content": "<|USER_TOKEN|>",
64
+ "lstrip": false,
65
+ "normalized": false,
66
+ "rstrip": false,
67
+ "single_word": false,
68
+ "special": false
69
+ },
70
+ "255003": {
71
+ "content": "<|CHATBOT_TOKEN|>",
72
+ "lstrip": false,
73
+ "normalized": false,
74
+ "rstrip": false,
75
+ "single_word": false,
76
+ "special": false
77
+ },
78
+ "255004": {
79
+ "content": "<|SYSTEM_TOKEN|>",
80
+ "lstrip": false,
81
+ "normalized": false,
82
+ "rstrip": false,
83
+ "single_word": false,
84
+ "special": false
85
+ },
86
+ "255005": {
87
+ "content": "<|NEW_FILE|>",
88
+ "lstrip": false,
89
+ "normalized": false,
90
+ "rstrip": false,
91
+ "single_word": false,
92
+ "special": true
93
+ },
94
+ "255006": {
95
+ "content": "<|BEGINNING_OF_PREFIX_FIM_TOKEN|>",
96
+ "lstrip": false,
97
+ "normalized": false,
98
+ "rstrip": false,
99
+ "single_word": false,
100
+ "special": true
101
+ },
102
+ "255007": {
103
+ "content": "<|BEGINNING_OF_MIDDLE_FIM_TOKEN|>",
104
+ "lstrip": false,
105
+ "normalized": false,
106
+ "rstrip": false,
107
+ "single_word": false,
108
+ "special": true
109
+ },
110
+ "255008": {
111
+ "content": "<|BEGINNING_OF_SUFFIX_FIM_TOKEN|>",
112
+ "lstrip": false,
113
+ "normalized": false,
114
+ "rstrip": false,
115
+ "single_word": false,
116
+ "special": true
117
+ },
118
+ "255009": {
119
+ "content": "<|END_OF_MIDDLE_FIM_TOKEN|>",
120
+ "lstrip": false,
121
+ "normalized": false,
122
+ "rstrip": false,
123
+ "single_word": false,
124
+ "special": true
125
+ },
126
+ "255010": {
127
+ "content": "<|START_THINKING|>",
128
+ "lstrip": false,
129
+ "normalized": false,
130
+ "rstrip": false,
131
+ "single_word": false,
132
+ "special": false
133
+ },
134
+ "255011": {
135
+ "content": "<|END_THINKING|>",
136
+ "lstrip": false,
137
+ "normalized": false,
138
+ "rstrip": false,
139
+ "single_word": false,
140
+ "special": false
141
+ },
142
+ "255012": {
143
+ "content": "<|START_RESPONSE|>",
144
+ "lstrip": false,
145
+ "normalized": false,
146
+ "rstrip": false,
147
+ "single_word": false,
148
+ "special": false
149
+ },
150
+ "255013": {
151
+ "content": "<|END_RESPONSE|>",
152
+ "lstrip": false,
153
+ "normalized": false,
154
+ "rstrip": false,
155
+ "single_word": false,
156
+ "special": false
157
+ },
158
+ "255014": {
159
+ "content": "<|START_ACTION|>",
160
+ "lstrip": false,
161
+ "normalized": false,
162
+ "rstrip": false,
163
+ "single_word": false,
164
+ "special": false
165
+ },
166
+ "255015": {
167
+ "content": "<|END_ACTION|>",
168
+ "lstrip": false,
169
+ "normalized": false,
170
+ "rstrip": false,
171
+ "single_word": false,
172
+ "special": false
173
+ },
174
+ "255016": {
175
+ "content": "<|START_TOOL_RESULT|>",
176
+ "lstrip": false,
177
+ "normalized": false,
178
+ "rstrip": false,
179
+ "single_word": false,
180
+ "special": false
181
+ },
182
+ "255017": {
183
+ "content": "<|END_TOOL_RESULT|>",
184
+ "lstrip": false,
185
+ "normalized": false,
186
+ "rstrip": false,
187
+ "single_word": false,
188
+ "special": false
189
+ },
190
+ "255018": {
191
+ "content": "<|USER_0_TOKEN|>",
192
+ "lstrip": false,
193
+ "normalized": false,
194
+ "rstrip": false,
195
+ "single_word": false,
196
+ "special": false
197
+ },
198
+ "255019": {
199
+ "content": "<|USER_1_TOKEN|>",
200
+ "lstrip": false,
201
+ "normalized": false,
202
+ "rstrip": false,
203
+ "single_word": false,
204
+ "special": false
205
+ },
206
+ "255020": {
207
+ "content": "<|USER_2_TOKEN|>",
208
+ "lstrip": false,
209
+ "normalized": false,
210
+ "rstrip": false,
211
+ "single_word": false,
212
+ "special": false
213
+ },
214
+ "255021": {
215
+ "content": "<|USER_3_TOKEN|>",
216
+ "lstrip": false,
217
+ "normalized": false,
218
+ "rstrip": false,
219
+ "single_word": false,
220
+ "special": false
221
+ },
222
+ "255022": {
223
+ "content": "<|USER_4_TOKEN|>",
224
+ "lstrip": false,
225
+ "normalized": false,
226
+ "rstrip": false,
227
+ "single_word": false,
228
+ "special": false
229
+ },
230
+ "255023": {
231
+ "content": "<|USER_5_TOKEN|>",
232
+ "lstrip": false,
233
+ "normalized": false,
234
+ "rstrip": false,
235
+ "single_word": false,
236
+ "special": false
237
+ },
238
+ "255024": {
239
+ "content": "<|USER_6_TOKEN|>",
240
+ "lstrip": false,
241
+ "normalized": false,
242
+ "rstrip": false,
243
+ "single_word": false,
244
+ "special": false
245
+ },
246
+ "255025": {
247
+ "content": "<|USER_7_TOKEN|>",
248
+ "lstrip": false,
249
+ "normalized": false,
250
+ "rstrip": false,
251
+ "single_word": false,
252
+ "special": false
253
+ },
254
+ "255026": {
255
+ "content": "<|USER_8_TOKEN|>",
256
+ "lstrip": false,
257
+ "normalized": false,
258
+ "rstrip": false,
259
+ "single_word": false,
260
+ "special": false
261
+ },
262
+ "255027": {
263
+ "content": "<|USER_9_TOKEN|>",
264
+ "lstrip": false,
265
+ "normalized": false,
266
+ "rstrip": false,
267
+ "single_word": false,
268
+ "special": false
269
+ },
270
+ "255028": {
271
+ "content": "<|VISION_START|>",
272
+ "lstrip": false,
273
+ "normalized": false,
274
+ "rstrip": false,
275
+ "single_word": false,
276
+ "special": true
277
+ },
278
+ "255029": {
279
+ "content": "<|VISION_END|>",
280
+ "lstrip": false,
281
+ "normalized": false,
282
+ "rstrip": false,
283
+ "single_word": false,
284
+ "special": true
285
+ },
286
+ "255030": {
287
+ "content": "<|VISION_PAD|>",
288
+ "lstrip": false,
289
+ "normalized": false,
290
+ "rstrip": false,
291
+ "single_word": false,
292
+ "special": true
293
+ },
294
+ "255031": {
295
+ "content": "<|IMAGE_PAD|>",
296
+ "lstrip": false,
297
+ "normalized": false,
298
+ "rstrip": false,
299
+ "single_word": false,
300
+ "special": true
301
+ },
302
+ "255032": {
303
+ "content": "<|VIDEO_PAD|>",
304
+ "lstrip": false,
305
+ "normalized": false,
306
+ "rstrip": false,
307
+ "single_word": false,
308
+ "special": true
309
+ }
310
+ },
311
+ "additional_special_tokens": [
312
+ "<|VISION_START|>",
313
+ "<|IMAGE_PAD|>",
314
+ "<|VISION_END|>",
315
+ "<|VISION_PAD|>",
316
+ "<|VIDEO_PAD|>"
317
+ ],
318
  "bos_token": "<BOS_TOKEN>",
319
  "clean_up_tokenization_spaces": false,
 
320
  "eos_token": "<|END_OF_TURN_TOKEN|>",
321
+ "extra_special_tokens": {},
322
  "image_token": "<|IMAGE_PAD|>",
 
323
  "legacy": true,
 
 
324
  "max_pixels": 3868706,
325
+ "merges_file": null,
326
  "min_pixels": 16384,
327
  "model_max_length": 1000000000000000019884624838656,
 
 
 
 
 
 
328
  "pad_token": "<PAD>",
329
  "padding_side": "right",
330
  "processor_class": "CohereCompassProcessor",
 
331
  "sp_model_kwargs": {},
332
  "spaces_between_special_tokens": false,
333
  "tokenizer_class": "CohereTokenizer",
 
335
  "use_default_system_prompt": false,
336
  "video_token": "<|VIDEO_PAD|>",
337
  "vision_end_token": "<|VISION_END|>",
338
+ "vision_start_token": "<|VISION_START|>",
339
+ "vocab_file": null
340
  }
video_preprocessor_config.json ADDED
@@ -0,0 +1,21 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "size": {
3
+ "longest_edge": 25165824,
4
+ "shortest_edge": 4096
5
+ },
6
+ "patch_size": 16,
7
+ "temporal_patch_size": 2,
8
+ "merge_size": 2,
9
+ "image_mean": [
10
+ 0.5,
11
+ 0.5,
12
+ 0.5
13
+ ],
14
+ "image_std": [
15
+ 0.5,
16
+ 0.5,
17
+ 0.5
18
+ ],
19
+ "processor_class": "CohereCompassProcessor",
20
+ "video_processor_type": "CohereCompassVideoProcessor"
21
+ }