This plugin honestly looks great, but I'm a bit reluctant when it comes to installing it because I'm concerned that I'll see a decrease in performance with it (since sometimes, the quality of the response of a model can be linked to the number of tokens it produces).
Would anyone be down to run benchmarks with this skill and see if we see and decline in performance or not?
Caveman does it: https://github.com/juliusbrussee/caveman
This plugin honestly looks great, but I'm a bit reluctant when it comes to installing it because I'm concerned that I'll see a decrease in performance with it (since sometimes, the quality of the response of a model can be linked to the number of tokens it produces).
Would anyone be down to run benchmarks with this skill and see if we see and decline in performance or not?
Caveman does it: https://github.com/juliusbrussee/caveman