ORIGINAL REDDIT POST
Treasure Hunt benchmarks available?
Using LLM’s for coding is legendary. Using it to try and solve treasure hunts however, absolute dogshit. Are there any benchmarks yet for this? It seems that if you can decipher vague clues, chain them together, prevent red herrings and such would really help…
Using LLM’s for coding is legendary. Using it to try and solve treasure hunts however, absolute dogshit. Are there any benchmarks yet for this? It seems that if you can decipher vague clues, chain them together, prevent red herrings and such would really help with general capability and adaptability.
Collected discussion
Hi! I have a website for bechmarking open-weight models. If you want to create any like you described, pm.
Hi, is the website public, and if so can you link it? I would be interested in independent benchmarks of open models