REDDIT 原始帖子

Treasure Hunt benchmarks available?

Using LLM’s for coding is legendary. Using it to try and solve treasure hunts however, absolute dogshit. Are there any benchmarks yet for this? It seems that if you can decipher vague clues, chain them together, prevent red herrings and such would really help…

原帖正文r/LocalLLaMA

Using LLM’s for coding is legendary. Using it to try and solve treasure hunts however, absolute dogshit. Are there any benchmarks yet for this? It seems that if you can decipher vague clues, chain them together, prevent red herrings and such would really help with general capability and adaptability.

已收录讨论

2 条评论

u/Equivalent_Job_2257

Hi! I have a website for bechmarking open-weight models. If you want to create any like you described, pm.

u/Fisent

Hi, is the website public, and if so can you link it? I would be interested in independent benchmarks of open models