GPT-6 Sol·Luna의 이미지 인식을 떨어뜨리던 버그가 고쳐졌다
OpenAI가 GPT-6 Sol과 Luna의 이미지 이해 성능을 떨어뜨리던 버그를 고쳤다고 밝혔고, API와 Codex의 시각 작업과 컴퓨터 사용 기능에서 결과가 나아진다고 했다. 벤치마크를 다시 돌린 Andon Labs에 따르면 Sol은 조금 올랐다. Luna는 크게 뛰었다. 일반 대화 평가는 좋지 않다. 사람이 블라인드로 고르는 Text Arena에서 GPT-6 Sol은 60위 근처로, 이전 GPT-5.6 Sol의 19위보다 크게 밀렸다는 지적이 나왔다. OpenAI DevDay는 미국 시간 29일이다.
출처 4건 보기· @OpenAIDevs, @andonlabs, @TokenGremlin 외 1
- @OpenAIDevsWe’ve fixed a bug that was degrading image understanding in GPT-6 Sol and GPT-6 Luna. You should now see better results on visual tasks in the API and Codex, including computer use.X ♥3.5천
- @andonlabsOpenAI had a bug, so we reran Blueprint-Bench for GPT-6 Sol and Luna. Sol improved slightly; Luna made a massive leap!X ♥359
- @TokenGremlinOne detail a lot of people missed: GPT-6 Sol is currently performing much worse than GPT-5.6 Sol in Text Arena. GPT-5.6 Sol sits around rank #19. GPT-6 Sol is around rank #60. That’s a massive drop in blind human preference tests for text. GPT-6 Sol may be cheaper and much stronger at coding and agentic work, but for normal ChatGPT-style conversations, this looks like a pretty brutal regression.X ♥1.2천
- @OpenAIDevsOpenAI DevDay roll call Where are you joining us from, and what are you building?X ♥3.4천