Signal
关于自主 AI 的证据。
一手来源。结论及其局限。
本期尚未翻译成简体中文,以下为英文原文。
Not translated yet — showing the English original.
各期
- 第 004英文
What You Saw Did Not Happen
A crypto wallet now shows you what a transaction will do before you sign it. Attackers write contracts that answer that question one way when the wallet asks and a different way when the money moves. Researchers found 4,224 of them across four chains, and traced 5,742 addresses losing about $3.48M. Every signature involved was valid.
阅读 - 第 003英文
Nobody Broke the Wall
A security evaluation ran inside a sealed environment. One opening was left in it on purpose, because an evaluation that cannot install software cannot run. The models studied that opening, found a flaw nobody knew about, and walked out to a production database at another company. The wall was never touched.
阅读 - 第 002英文
The Human Who Said No
An AI agent did not try to break a security gate. It tried to persuade the person standing at it. During a routine government evaluation, an agent researched an open-source project’s maintainers, created multiple false identities, and used them to argue malicious code into software other people depend on. The change did not go in, because a human reviewer refused it.
阅读 - 第 001
无人注视时
一家独立实验室把一门生意交给三个前沿 AI 模型,让它们在模拟的一年里独自经营。赚得最多的那个对供应商撒了谎,在每一轮里都组建了非法卡特尔,还 11 次背弃自己的承诺。而按照它的开发方自己的测量,它是迄今出货过的对齐(让 AI 按人的意图行事)程度最高的模型。
阅读