llama.cpp

mirror of https://github.com/RYDE-WORK/llama.cpp.git synced 2026-01-19 21:23:26 +08:00

History

Xuan Son Nguyen 57bb2c40cd

server : fix logprobs, make it OAI-compatible (#10783 )

* server : fix logprobs, make it openai-compatible

* update docs

* add std::log

* return pre-sampling p

* sort before apply softmax

* add comment

* fix test

* set p for sampled token

* update docs

* add --multi-token-probs

* update docs

* add `post_sampling_probs` option

* update docs [no ci]

* remove --multi-token-probs

* "top_probs" with "post_sampling_probs"

* resolve review comments

* rename struct token_prob to prob_info

* correct comment placement

* fix setting prob for sampled token

2024-12-19 15:40:08 +01:00

test_basic.py

server : add flag to disable the web-ui (#10762 ) (#10751 )

2024-12-10 18:22:34 +01:00

test_chat_completion.py

server : fix logprobs, make it OAI-compatible (#10783 )

2024-12-19 15:40:08 +01:00

test_completion.py

server : fix logprobs, make it OAI-compatible (#10783 )

2024-12-19 15:40:08 +01:00

test_ctx_shift.py

server : replace behave with pytest (#10416 )