Issue Type
Outdated content
Document Location
README.md, https://open-codereview.ai/benchmark
Description
- Benchmarks are using old models: the top benchmark is still for 4.6-Opus. We're on 5.5 now.
- Curious about Fable, although totally understand hesitancy on spending that much.
- In real life, our team is also looking at smaller, cheaper models. Curious about new Sonnet 5.5, but right now in production our team is using Luna 6 for efficiency. Benchmarks should include those as well!
Rather than running it all ourselves, we should have the maintainers run these benchmarks for efficiency.
Suggested Change
No response
Issue Type
Outdated content
Document Location
README.md, https://open-codereview.ai/benchmark
Description
Rather than running it all ourselves, we should have the maintainers run these benchmarks for efficiency.
Suggested Change
No response