From 6bcec86fa025c6a503016684eb39cedb89e42afb Mon Sep 17 00:00:00 2001 From: Azure Date: Thu, 15 Aug 2024 16:54:31 +0000 Subject: [PATCH] update README --- README.md | 5 +++++ 1 file changed, 5 insertions(+) diff --git a/README.md b/README.md index 8c5f505..b70a59a 100644 --- a/README.md +++ b/README.md @@ -22,6 +22,11 @@ interface, RESTful APIs compliant with OpenAI and Ollama, and even a simplified

Our vision for KTransformers is to serve as a flexible platform for experimenting with innovative LLM inference optimizations. Please let us know if you need any other features. +

🔥 Updates

+ +* **Aug 14, 2024**: Support llamfile as linear backend, +* **Aug 12, 2024**: Support multiple GPU; Support new model: mixtral 8\*7B and 8\*22B; Support q2k, q3k, q5k dequant on gpu. +* **Aug 9, 2024**: Support windows native.

🔥 Show Cases

GPT-4-level Local VSCode Copilot on a Desktop with only 24GB VRAM