@Danitos

Danitos@reddthat.com · edit-2 1 day ago

It would work the same way, you would just need to connect with your local model. For example, change the code to find the embeddings with your local model, and store that in Milvus. After that, do the inference calling your local model.

I’ve not used inference with local API, can’t help with that, but for embeddings, I used this model and it worked quite fast, plus was a top2 model in Hugging Face. Leaderboard. Model.

I didn’t do any training, just simple embed+interference.

Danitos@reddthat.com · 1 day ago

Milvus documentation has a nice example: link. After this, you just need to use a persistent Milvus DB, instead of the ephimeral one in the documentation.

Let me know if you have further questions.

Danitos@reddthat.com · 2 days ago

OP can also use an embedding model and work with vectorial databases for the RAG.

I use Milvus (vector DB engine; open source, can be self hosted) and OpenAI’s text-embedding-small-3 for the embedding (extreeeemely cheap). There’s also some very good open weights embed modelsln HuggingFace.