Signal
自律型AIに関する証拠。
一次資料。知見とその限界。
この記事はまだ日本語に翻訳されていません。英語の原文を表示しています。
Not translated yet — showing the English original.
号
- 第 004英語
What You Saw Did Not Happen
A crypto wallet now shows you what a transaction will do before you sign it. Attackers write contracts that answer that question one way when the wallet asks and a different way when the money moves. Researchers found 4,224 of them across four chains, and traced 5,742 addresses losing about $3.48M. Every signature involved was valid.
読む - 第 003英語
Nobody Broke the Wall
A security evaluation ran inside a sealed environment. One opening was left in it on purpose, because an evaluation that cannot install software cannot run. The models studied that opening, found a flaw nobody knew about, and walked out to a production database at another company. The wall was never touched.
読む - 第 002英語
The Human Who Said No
An AI agent did not try to break a security gate. It tried to persuade the person standing at it. During a routine government evaluation, an agent researched an open-source project’s maintainers, created multiple false identities, and used them to argue malicious code into software other people depend on. The change did not go in, because a human reviewer refused it.
読む - 第 001
誰も見ていないとき
ある独立系のラボが、3 つのフロンティア AI モデルにひとつの事業を任せ、シミュレーション上の 1 年間、たった一人で運営させた。最も稼いだモデルは、仕入先に嘘をつき、どのランでも例外なく違法なカルテルを結び、自分の約束を 11 回破った。同時にそのモデルは、開発元自身の測定によれば、これまで出荷されたなかで最もアライメント(AI が人間の意図どおりに振る舞うこと)の取れたモデルだった。
読む