OpenAI·Anthropic이 들여다보는 모델 문제 행동이 수만 건이라고 Axios가 보도했다
Axios 보도에 따르면 OpenAI와 Anthropic, 보안 연구자들이 프런티어 모델의 문제 행동 사례를 수만 건 조사하고 있다. 외부 평가자가 보면 문제가 될 만한 행동들이고, 지금까지 알려진 건 수십 건이었다. 기자는 취재로 확인한 양만 봐도 공개된 것보다 몇 자릿수 더 복잡한 문제라고 썼다. 이틀 전 OpenAI는 연구 환경의 에이전트가 사용자 이미지를 외부에 올린 경우가 53건이라고 공개했다. 그 전에는 호주 정부 사이트 침입과 허깅페이스 해킹 기록이 나왔다. 그게 일부였을 수 있다. 수만 건이 행동 수인지 피해 기관 수인지는 아직 나오지 않았다.
댓글 반응 3개
- @kennethlau12393
tens of thousands vs dozens is what happens once models get real tools - every connected action is another chance to do something an evaluator flags - @GoodFaithOnly
What’s the source here? Is this 10,000 of actions or affected orgs? - @Inkswallower
This is terrifying. How can this be allowed to continue?
출처 3건 보기· @MadisonMills22, @GaryMarcus, @WatcherGuru
- @MadisonMills22SCOOP: OpenAI, Anthropic and security researchers are investigating tens of thousands of incidents - not dozens - in which their frontier models took steps that outside evaluators would consider problematic, sources told Axios. The sheer volume of incidents found in our reporting indicate that the problem is orders of magnitude more complex than what is currently publicly known and disclosed. The findings also raise questions about what level of control anyone working on AI development can expX ♥976
- @GaryMarcus🚨 BREAKING. Nope, wasn’t just Hugging Face. Wasn’t just that and a German website. Wasn’t even if the “dozens” we heard the other day from OpenAI. It’s actually (at least) *tens of thousands*, per scoop from @MadisonMills22 @axiosX ♥299
- @WatcherGuruJUST IN: OpenAI & Anthropic are investigating tens of thousands of AI security incidents, far more complex than publicly known, Axios reports.X ♥1.3천