Next.js 평가에서 세 모델이 97%로 동률 — 남은 차이는 가격이었다
Vercel의 기예르모 라우시가 새로 돌린 Next.js 평가 결과를 올렸다. Opus 5.5, GPT-6 Sol, Fable 5.1이 모두 97%로 같았고 Grok 4.7이 94%로 뒤를 이었다. 라우시는 Grok이 두 배에서 일곱 배 싸다는 점을 따로 짚었다. 일론 머스크도 Grok 4.7이 작은 모델치고 꽤 잘한다고 거들었다. 순위를 두고는 말이 갈린다. 한쪽에서는 Opus 5.5, GPT-6 Astra, Fable 5.1, GPT-6 Sol 순으로 줄을 세웠다. 반대로 OpenAI가 낸 Sol·Luna 벤치마크 표는 역대 가장 읽기 어렵다는 불평도 나왔다. 평가 점수가 한곳으로 모이면서, 비교의 무게가 점수에서 작업당 비용으로 옮겨가고 있다.
댓글 반응 3개
- @trekedge
GPT 6 Sol and Opus have saturated this benchmark. Time for a harder one. - @cloneisjun
how many runs per model is that 97% averaged over, since the gap to grok is only 3 points - @shahab_jozdani
I'd pay the >2x costs if that 3% means it saves more time and helps get a job done faster, which is apparently the case at the moment.
출처 4건 보기· @rauchg, @bridgemindai, @alexgetmancom 외 1
- @rauchgWe ran fresh Next.js evals. The tally: ① Opus 5.5 [𝟿𝟽%] ② GPT 6 Sol [𝟿𝟽%] ③ Fable 5.1 [𝟿𝟽%] ④ Grok 4.7 [𝟿𝟺%] Notably, Grok is 2x-7x cheaperX ♥297
- @bridgemindaiClaude Opus 5.5 > GPT 6 Astra > Fable 5.1 > GPT 6 SolX ♥298
- @alexgetmancomGPT-6 Sol and Luna benchmarks are out. These are the worst benchmarks in OpenAI history. They’re basically impossible to read. If anyone can make sense of them, let me know.X ♥189
- @elonmuskInteresting. Grok 4.7 is performing fairly well for a smallish model.X ♥2.3천