ORIGINAL REDDIT POST

Treasure Hunt benchmarks available?

Using LLM’s for coding is legendary. Using it to try and solve treasure hunts however, absolute dogshit. Are there any benchmarks yet for this? It seems that if you can decipher vague clues, chain them together, prevent red herrings and such would really help…

Original postr/LocalLLaMA

Using LLM’s for coding is legendary. Using it to try and solve treasure hunts however, absolute dogshit. Are there any benchmarks yet for this? It seems that if you can decipher vague clues, chain them together, prevent red herrings and such would really help with general capability and adaptability.

Collected discussion

2 comments

u/Equivalent_Job_2257

Hi! I have a website for bechmarking open-weight models. If you want to create any like you described, pm.

u/Fisent

Hi, is the website public, and if so can you link it? I would be interested in independent benchmarks of open models

Treasure Hunt benchmarks available?