94 Commits

Author SHA1 Message Date
Azure
ff6b265e53 Mock triton mla due to precision issue 2025-02-16 06:03:12 +00:00
Atream
c5f036e8a4
Merge pull request #333 from kvcache-ai/feat_experts_gpu
toy support for experts on GPU, no CUDA Graph
2025-02-15 23:30:24 +08:00
Atream
c189d55bd1 toy support for experts on GPU, no CUDA Graph 2025-02-15 15:16:00 +00:00
wang jiahao
ae8da019c1
Merge pull request #330 from hrz6976/fix-nonetype
thanks..., I was about to submit and found that you had already modified it. Thank you for your contribution
2025-02-15 22:50:51 +08:00
Azure
56382aa869
Merge pull request #290 from ZhangShuaiyi/dev/check_gguf_path
ensure that gguf_path argument is a directory.
2025-02-15 22:48:57 +08:00
12f23eddde
4516282ccc Fix NoneType object has no attribute zero_ 2025-02-15 22:04:45 +08:00
UnicornChan
65d73ea3f8
Merge pull request #317 from kvcache-ai/develop-0.2.1
[feature] update  docker image and entrypoint
2025-02-15 16:19:18 +08:00
chenxl
0e4b7a3929 [feature] update docker image and entrypoint 2025-02-15 07:55:33 +00:00
Atream
92399283b6
Update attention.py 2025-02-15 15:43:35 +08:00
Atream
d90749d35d
Update triton_attention.py 2025-02-15 15:41:01 +08:00
Shuaiyi
22280bf17f get dirname if gguf_path is a file 2025-02-15 07:08:22 +00:00
Atream
1946493f2d warm_up before capture 2025-02-14 15:52:21 +00:00
Atream
885a91e7db
Merge pull request #294 from kvcache-ai/feat-fast-MLA
Feat fast mla
2025-02-14 19:40:36 +08:00
Atream
1084d4e4b4 linux support triton MLA kernel 2025-02-14 11:38:55 +00:00
Azure
95c81eaf01
Merge branch 'kvcache-ai:main' into main 2025-02-14 11:53:52 +08:00
Azure
b7653b9c4f add V3/R1 8 gpu yaml example 2025-02-14 02:56:13 +00:00
Azure
ae5d9e11a9
Merge pull request #227 from hrz6976/main
Add a lock to server inference()
2025-02-14 10:35:11 +08:00
Atream
bb35dc5b0d init support for MLA using Attention kernel 2025-02-13 15:01:14 +00:00
hrz6976
2c3dcd9774 Add a lock to server inference() 2025-02-13 10:05:22 +00:00
ZiWei Yuan
76b081879a
Merge pull request #224 from kvcache-ai/server_support
Server support
2025-02-13 17:28:08 +08:00
liam
8d5ebe49ab 📝 fix some debug output and update doc 2025-02-13 17:25:12 +08:00
liam
c74453d8ca 📝 add doc support and fix bug in qwen2 2025-02-13 16:37:43 +08:00
MorphisZhang
aea4243712 Add optimization config for Deepseek V3/R1 with 4 GPUs 2025-02-13 16:32:28 +08:00
Azure
101db0e9de Merge branch 'main' into update-yaml 2025-02-12 08:56:03 +00:00
Azure
3897f001f5 update FAQ 2025-02-12 08:50:58 +00:00
ZiWei Yuan
4ae2e81c38
Merge pull request #152 from kvcache-ai/server_support
Server support
2025-02-12 12:45:13 +08:00
liam
4385e85096 support force thinking 2025-02-12 12:43:53 +08:00
Azure
0564ac8465 update marlin expert example 2025-02-12 04:11:00 +00:00
liam
6f3a39be08 update force_think config 2025-02-12 12:10:16 +08:00
liam
e536e1420d update force_think 2025-02-12 11:42:55 +08:00
liam
d07087a7e2 support R1 force thinking 2025-02-11 15:43:41 +08:00
UnicornChan
7527619f53
Merge pull request #122 from kvcache-ai/feat-DeepSeekV3
[Feat] add support to DeepSeekV3
2025-02-10 13:54:46 +08:00
liam
83401dbb3b ready to publish 2025-02-10 12:29:23 +08:00
unicornchan
c7e6d09068 [feature] update version and github action jobs for package 2025-02-10 01:00:57 +00:00
liam
098602b08f v0.2 ongoing 2025-02-09 22:41:14 +08:00
liam
bf1d413be0 Merge branch 'feat-DeepSeekV3' of github.com:kvcache-ai/ktransformers into feat-DeepSeekV3 2025-02-08 13:17:10 +08:00
liam
c18ecd7b7f add flush print in local_chat output and change default optimize yaml of deepseekv3 to single gpu 2025-02-08 13:15:52 +08:00
RodriMora
b1bff2a405 Added simple /models endpoint to work with frontends that don't allow bypass check like Openweb-ui 2025-02-07 10:30:39 +01:00
Azure
c4d9bc6670 support KExpertsMarlin backend 2025-02-07 05:57:40 +00:00
liam
0262f954c7 Merge branch 'feat-DeepSeekV3' of github.com:kvcache-ai/ktransformers into feat-DeepSeekV3 2025-02-06 22:41:25 +08:00
liam
3dca28d23b fix moe.cpp int overflow problem 2025-02-06 22:39:16 +08:00
Azure
027b11266c modify moeinfer param 2025-02-06 14:07:38 +00:00
Azure
ee24a27001 update v3 single gpu rule yaml; 2025-02-04 16:14:35 +00:00
Azure
907251c743 done support deepseekv3 2025-02-04 15:53:38 +00:00
Azure
f748cd29f0 fix rope; update moegate 2025-02-01 18:05:45 +00:00
Azure
f873558a89 update rope calculation; update modeling.py; update gate for moe 2025-02-01 07:32:21 +00:00
Azure
5a50b34627 fix hard coding caused by rope dim calculation, load from config now 2025-01-31 15:25:50 +00:00
Azure
476b1d8dc6 support deepseekv3; runable but have precition problem 2025-01-31 08:27:24 +00:00
liam
04cebec4bb rm opt config path default value and fix some config logic bug 2024-11-14 20:02:30 +08:00
liam
c2b4dc805c 🚑️:roll back transformer.py and find that it's multiple chat hsitory have minor accurate error 2024-11-04 14:02:19 +08:00