買幣
行情
現貨
合約
理財
活動
更多
reward-center新手專區
信息首頁快訊詳情
Anthropic provides Claude with on-chain lock: changing a single sentence in the context invalidates the entire model distillation.
  • OPUS0%

Perceive Beating AI News: Anthropic has started adding a "Context Lock" to Claude's encrypted musings. Fable 5.1 mandates that any encrypted musing generated via API must be returned exactly as it was generated, along with the system prompts, tools, and historical messages used during its creation. Any modification to the previous content will trigger an API error or result in the rejection of the musing.

This change is aimed at preventing model distillation. Previously, researchers discovered that although Claude's encrypted musings were incomprehensible to the user, they could be fed back to Anthropic's interoperable model for interpretation. An attacker could first have Opus produce high-quality reasoning, then insert the encrypted block to a less secure safeguard like Haiku, prompting it to articulate Opus's complete musings. This not only involves copying the large model's answers but also taking away the entire problem-solving process written on the scratch paper.

Anthropic had already thwarted one round of attacks when it released Fable 5 in June by binding the encrypted musings to the model, preventing them from being handed to smaller models like Haiku for decryption. Fable 5.1 has now introduced "Dialog Binding": if there are any modifications to the preceding prompts, tools, or chat records, the old musings are invalidated.

Previously, the defense was against "model swapping for musing theft"; now, even "swapping context to sleuth musings" has been addressed.

來源:BlockBeats

免責聲明:當前內容均來自第三方觀點或由AI直接翻譯第三方觀點,CoinEx不保證內容的真實性、準確性和原創性,不構成CoinEx相關的任何投資建議。數字資產價格波動劇烈,請注意潛在風險。

熱搜榜
  • 幣種
    價格
    24H漲跌