Azure
|
ff6b265e53
|
Mock triton mla due to precision issue
|
2025-02-16 06:03:12 +00:00 |
|
Atream
|
c5f036e8a4
|
Merge pull request #333 from kvcache-ai/feat_experts_gpu
toy support for experts on GPU, no CUDA Graph
|
2025-02-15 23:30:24 +08:00 |
|
Atream
|
c189d55bd1
|
toy support for experts on GPU, no CUDA Graph
|
2025-02-15 15:16:00 +00:00 |
|
wang jiahao
|
ae8da019c1
|
Merge pull request #330 from hrz6976/fix-nonetype
thanks..., I was about to submit and found that you had already modified it. Thank you for your contribution
|
2025-02-15 22:50:51 +08:00 |
|
Azure
|
56382aa869
|
Merge pull request #290 from ZhangShuaiyi/dev/check_gguf_path
ensure that gguf_path argument is a directory.
|
2025-02-15 22:48:57 +08:00 |
|
12f23eddde
|
4516282ccc
|
Fix NoneType object has no attribute zero_
|
2025-02-15 22:04:45 +08:00 |
|
UnicornChan
|
65d73ea3f8
|
Merge pull request #317 from kvcache-ai/develop-0.2.1
[feature] update docker image and entrypoint
|
2025-02-15 16:19:18 +08:00 |
|
chenxl
|
0e4b7a3929
|
[feature] update docker image and entrypoint
|
2025-02-15 07:55:33 +00:00 |
|
Atream
|
92399283b6
|
Update attention.py
|
2025-02-15 15:43:35 +08:00 |
|
Atream
|
d90749d35d
|
Update triton_attention.py
|
2025-02-15 15:41:01 +08:00 |
|
Shuaiyi
|
22280bf17f
|
get dirname if gguf_path is a file
|
2025-02-15 07:08:22 +00:00 |
|
Atream
|
1946493f2d
|
warm_up before capture
|
2025-02-14 15:52:21 +00:00 |
|
Atream
|
885a91e7db
|
Merge pull request #294 from kvcache-ai/feat-fast-MLA
Feat fast mla
|
2025-02-14 19:40:36 +08:00 |
|
Atream
|
1084d4e4b4
|
linux support triton MLA kernel
|
2025-02-14 11:38:55 +00:00 |
|
Azure
|
95c81eaf01
|
Merge branch 'kvcache-ai:main' into main
|
2025-02-14 11:53:52 +08:00 |
|
Azure
|
b7653b9c4f
|
add V3/R1 8 gpu yaml example
|
2025-02-14 02:56:13 +00:00 |
|
Azure
|
ae5d9e11a9
|
Merge pull request #227 from hrz6976/main
Add a lock to server inference()
|
2025-02-14 10:35:11 +08:00 |
|
Atream
|
bb35dc5b0d
|
init support for MLA using Attention kernel
|
2025-02-13 15:01:14 +00:00 |
|
hrz6976
|
2c3dcd9774
|
Add a lock to server inference()
|
2025-02-13 10:05:22 +00:00 |
|
ZiWei Yuan
|
76b081879a
|
Merge pull request #224 from kvcache-ai/server_support
Server support
|
2025-02-13 17:28:08 +08:00 |
|
liam
|
8d5ebe49ab
|
📝 ⚡ fix some debug output and update doc
|
2025-02-13 17:25:12 +08:00 |
|
liam
|
c74453d8ca
|
📝 add doc support and fix bug in qwen2
|
2025-02-13 16:37:43 +08:00 |
|
MorphisZhang
|
aea4243712
|
Add optimization config for Deepseek V3/R1 with 4 GPUs
|
2025-02-13 16:32:28 +08:00 |
|
Azure
|
101db0e9de
|
Merge branch 'main' into update-yaml
|
2025-02-12 08:56:03 +00:00 |
|
Azure
|
3897f001f5
|
update FAQ
|
2025-02-12 08:50:58 +00:00 |
|
ZiWei Yuan
|
4ae2e81c38
|
Merge pull request #152 from kvcache-ai/server_support
Server support
|
2025-02-12 12:45:13 +08:00 |
|
liam
|
4385e85096
|
⚡ support force thinking
|
2025-02-12 12:43:53 +08:00 |
|
Azure
|
0564ac8465
|
update marlin expert example
|
2025-02-12 04:11:00 +00:00 |
|
liam
|
6f3a39be08
|
⚡ update force_think config
|
2025-02-12 12:10:16 +08:00 |
|
liam
|
e536e1420d
|
⚡ update force_think
|
2025-02-12 11:42:55 +08:00 |
|
liam
|
d07087a7e2
|
⚡ support R1 force thinking
|
2025-02-11 15:43:41 +08:00 |
|
UnicornChan
|
7527619f53
|
Merge pull request #122 from kvcache-ai/feat-DeepSeekV3
[Feat] add support to DeepSeekV3
|
2025-02-10 13:54:46 +08:00 |
|
liam
|
83401dbb3b
|
⚡ ready to publish
|
2025-02-10 12:29:23 +08:00 |
|
unicornchan
|
c7e6d09068
|
[feature] update version and github action jobs for package
|
2025-02-10 01:00:57 +00:00 |
|
liam
|
098602b08f
|
⚡ v0.2 ongoing
|
2025-02-09 22:41:14 +08:00 |
|
liam
|
bf1d413be0
|
Merge branch 'feat-DeepSeekV3' of github.com:kvcache-ai/ktransformers into feat-DeepSeekV3
|
2025-02-08 13:17:10 +08:00 |
|
liam
|
c18ecd7b7f
|
⚡ add flush print in local_chat output and change default optimize yaml of deepseekv3 to single gpu
|
2025-02-08 13:15:52 +08:00 |
|
RodriMora
|
b1bff2a405
|
Added simple /models endpoint to work with frontends that don't allow bypass check like Openweb-ui
|
2025-02-07 10:30:39 +01:00 |
|
Azure
|
c4d9bc6670
|
support KExpertsMarlin backend
|
2025-02-07 05:57:40 +00:00 |
|
liam
|
0262f954c7
|
Merge branch 'feat-DeepSeekV3' of github.com:kvcache-ai/ktransformers into feat-DeepSeekV3
|
2025-02-06 22:41:25 +08:00 |
|
liam
|
3dca28d23b
|
⚡ fix moe.cpp int overflow problem
|
2025-02-06 22:39:16 +08:00 |
|
Azure
|
027b11266c
|
modify moeinfer param
|
2025-02-06 14:07:38 +00:00 |
|
Azure
|
ee24a27001
|
update v3 single gpu rule yaml;
|
2025-02-04 16:14:35 +00:00 |
|
Azure
|
907251c743
|
done support deepseekv3
|
2025-02-04 15:53:38 +00:00 |
|
Azure
|
f748cd29f0
|
fix rope; update moegate
|
2025-02-01 18:05:45 +00:00 |
|
Azure
|
f873558a89
|
update rope calculation; update modeling.py; update gate for moe
|
2025-02-01 07:32:21 +00:00 |
|
Azure
|
5a50b34627
|
fix hard coding caused by rope dim calculation, load from config now
|
2025-01-31 15:25:50 +00:00 |
|
Azure
|
476b1d8dc6
|
support deepseekv3; runable but have precition problem
|
2025-01-31 08:27:24 +00:00 |
|
liam
|
04cebec4bb
|
⚡ rm opt config path default value and fix some config logic bug
|
2024-11-14 20:02:30 +08:00 |
|
liam
|
c2b4dc805c
|
🚑️:roll back transformer.py and find that it's multiple chat hsitory have minor accurate error
|
2024-11-04 14:02:19 +08:00 |
|