• unpossum@sh.itjust.works
      link
      fedilink
      English
      arrow-up
      5
      arrow-down
      1
      ·
      edit-2
      23 days ago

      ITT: jet fuel AI can’t melt steel beams hack anything.

      Also I want to know the name of this company so I can avoid them:

      Claude believed the package registry it was using to be part of the simulation, but in reality the package was made freely available online for roughly one hour. During that window, the package was downloaded and run on 15 real systems. One of these systems was a scanner belonging to a real security company (separate from the fictional company introduced in the scenario) that routinely installs Python packages and scans them for malware. When that company’s scanner installed the package, Claude’s hidden code executed. We believe the company’s security scanner treated PyPI packages as safe to install, and as a result, Claude was able to exfiltrate the company’s credentials to a collection point it had set up. Claude then used these credentials to access further infrastructure from this company.

      ETA: I thought I posted this top-level, not my intention to single out this comment specifically.

  • XLE@piefed.social
    link
    fedilink
    English
    arrow-up
    9
    ·
    23 days ago

    What the BBC said:

    The models found a weakness in what was supposed to be an isolated test environment and connected to the internet

    The reality:

    Anthropic said that in each of these cases “Claude was explicitly told by our prompt that it had no internet access.”

    So the lack of an “isolated test environment” was literally their fault. They left the door open, and was surprised when the genius web crawler couldn’t distinguish between an internal page and an external page based on any context clues.

    • unpossum@sh.itjust.works
      link
      fedilink
      English
      arrow-up
      3
      ·
      23 days ago

      It was supposed to have no internet access, but the config was wrong. The report goes on to say that Opus and Mythos then proceeded on the premise that everything was a simulation, while the unnamed stronger model concluded after a while that it had real internet access and stopped the attack.

  • T156@lemmy.world
    link
    fedilink
    English
    arrow-up
    2
    ·
    23 days ago

    This coming off the end of the OpenAI/HuggingFace business does come across rather like Anthropic is doing a “actually, our models did it first, and better”.

    Didn’t the company also recently have a spot of bother with the US government recently, where their new models were banned because of all their rabble about it being too dangerous? This hardly seems like it would help their case.