Update sample code links
Browse files
README.md
CHANGED
|
@@ -30,7 +30,7 @@ You can use CoPE-B-A4B in three ways:
|
|
| 30 |
|
| 31 |
- **[Zentropi API](https://zentropi.ai/api)** — fastest path, with a generous free tier (no infra required)
|
| 32 |
- **Self-hosted vLLM** — for production-scale serving on your own infrastructure
|
| 33 |
-
- **Direct inference in Python** — load via Transformers; see this [Colab notebook](https://colab.research.google.com/drive/
|
| 34 |
|
| 35 |
See the [Running the Model](#running-the-model) section below for details on each.
|
| 36 |
|
|
@@ -375,7 +375,7 @@ The easiest way to get started with this model is to use it through the [Zentrop
|
|
| 375 |
|
| 376 |
### via Direct Inference (Python)
|
| 377 |
|
| 378 |
-
To call the model directly via Transformers, see this runnable [Colab notebook](https://colab.research.google.com/drive/
|
| 379 |
|
| 380 |
### via Self-Hosting (vLLM)
|
| 381 |
|
|
@@ -393,7 +393,7 @@ If you're currently using CoPE-A-9B and moving to CoPE-B-A4B, three things to kn
|
|
| 393 |
|
| 394 |
### 1. CoPE-B uses the Gemma-4 chat template
|
| 395 |
|
| 396 |
-
CoPE-B's prompt must be passed through `apply_chat_template` as a user-turn message — the answer comes back as the assistant-turn output. If your CoPE-A code path raw-concatenates the prompt directly, that pattern will not work with CoPE-B. See the [Input Format](#input-format) section above or the runnable [Colab notebook](https://colab.research.google.com/drive/
|
| 397 |
|
| 398 |
Note also that the CoPE-B prompt is leaner than CoPE-A's: there is no `INSTRUCTIONS` header or `ANSWER` footer to include — the chat template's role markers replace them.
|
| 399 |
|
|
|
|
| 30 |
|
| 31 |
- **[Zentropi API](https://zentropi.ai/api)** — fastest path, with a generous free tier (no infra required)
|
| 32 |
- **Self-hosted vLLM** — for production-scale serving on your own infrastructure
|
| 33 |
+
- **Direct inference in Python** — load via Transformers; see this [Colab notebook](https://colab.research.google.com/drive/1JD8OIa3yZYfVbeY81ao03lrvg0aS-6SQ) for a working example
|
| 34 |
|
| 35 |
See the [Running the Model](#running-the-model) section below for details on each.
|
| 36 |
|
|
|
|
| 375 |
|
| 376 |
### via Direct Inference (Python)
|
| 377 |
|
| 378 |
+
To call the model directly via Transformers, see this runnable [Colab notebook](https://colab.research.google.com/drive/1JD8OIa3yZYfVbeY81ao03lrvg0aS-6SQ). It loads CoPE-B-A4B from the Hub in bf16, applies the proper prompt template, and shows a complete worked example end-to-end.
|
| 379 |
|
| 380 |
### via Self-Hosting (vLLM)
|
| 381 |
|
|
|
|
| 393 |
|
| 394 |
### 1. CoPE-B uses the Gemma-4 chat template
|
| 395 |
|
| 396 |
+
CoPE-B's prompt must be passed through `apply_chat_template` as a user-turn message — the answer comes back as the assistant-turn output. If your CoPE-A code path raw-concatenates the prompt directly, that pattern will not work with CoPE-B. See the [Input Format](#input-format) section above or the runnable [Colab notebook](https://colab.research.google.com/drive/1JD8OIa3yZYfVbeY81ao03lrvg0aS-6SQ) for the exact pattern.
|
| 397 |
|
| 398 |
Note also that the CoPE-B prompt is leaner than CoPE-A's: there is no `INSTRUCTIONS` header or `ANSWER` footer to include — the chat template's role markers replace them.
|
| 399 |
|