Keep these models coming! And hopefully with vision modality included in the future

#1
by biologin - opened

Great work with this model, really impressed with its capacity. Just missing the vision modality and it is literally the perfect sub-agent model.

Not sure how hard it would be to add a vision encoder of some sort (even if basic) to a model like this, but I think it would be a major addition. Most people realistically only have access to 8 GB VRAM, so that new model at Q8 would ideally occupy only around 3 GB VRAM and the rest would go for context (100k +).

Any future 2B/2.5B models with vision planned?

OpenBMB org

Thanks for the kind words, glad the model is working well for you!
On vision: multimodal is on our roadmap and we do plan to open-source those models β€” stay tuned.

Sign up or log in to comment