github.com/antirez/llam...
And a GGUF file you can use in order to run the inference with just 128 GB of RAM:
huggingface.co/antirez/deep...
github.com/antirez/llam...
And a GGUF file you can use in order to run the inference with just 128 GB of RAM:
huggingface.co/antirez/deep...
This project would have been impossible without the existence of llama.cpp and GGML and the work of Georgi Gerganov and all the other contributors. Thanks!
This project would have been impossible without the existence of llama.cpp and GGML and the work of Georgi Gerganov and all the other contributors. Thanks!
github.com/antirez/hist...
github.com/antirez/hist...
This is by far the most practical model for quick inference on low end CPUs. Perfect for things like transcribing bots, voice-driven UIs, and so forth.
This is by far the most practical model for quick inference on low end CPUs. Perfect for things like transcribing bots, voice-driven UIs, and so forth.
github.com/antirez/c64-...
github.com/antirez/c64-...