Evals - anyone have Claude results? #3070
Unanswered
leppikallio
asked this question in
Q&A
Replies: 0 comments
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Weird question, I know and the reason for the question (or rather why I am not running them myself to figure out) is longer story involving bitter feelings and salt poured in countles tiny paper cuts and... anyway, there's a reason to come up with the topic.
I wired the Gas Town with OpenCode & GPT models / model variants out of curiosity, full well knowing that GPT most probably won't behave nicely. So then I ran the evals and results do not exactly encourage. The failure modes & variety makes kind of confirms the "not-fit-for-the-purpose".
But, indeed out of curiosity, what's the Claude's track record with the current GT & evals?
(I can of course attach the "analysis" GPT created if someone is curious)
All reactions