Most mass scrapers, on the other hand, simply grab the raw HTML underneath. ShieldFont exploits this difference through an automated process called OpenType glyph substitution.

That said, because the whole defense rests on scrapers reading code rather than screens, taking a screenshot of a shielded page and running OCR on the image can still recover the real words.

Screen readers used by blind readers also work from the code, so they read the decoys aloud. ShieldFont ships with a beta feature that provides those readers with the real text instead.

    • FTonsilStones@lemmy.caOP
      link
      fedilink
      English
      arrow-up
      54
      arrow-down
      1
      ·
      23 days ago

      Per the article, yes:

      Screen readers used by blind readers also work from the code, so they read the decoys aloud.

      But:

      ShieldFont ships with a beta feature that provides those readers with the real text instead.

      • ViatorOmnium@piefed.social
        link
        fedilink
        English
        arrow-up
        58
        ·
        23 days ago

        ShieldFont ships with a beta feature that provides those readers with the real text instead.

        AI scrappers will just pretend to be screen readers then.

        And if the approach becomes popular they will just OCR the text instead.

        • cley_faye@lemmy.world
          link
          fedilink
          English
          arrow-up
          7
          ·
          23 days ago

          AI scrappers will just pretend to be screen readers then

          There’s a fair chance they’re already doing that. It provides better insight on the content, less formatting to handle, and even visual stuff gets text alternatives.

        • Sims@lemmy.ml
          link
          fedilink
          English
          arrow-up
          5
          ·
          23 days ago

          It is also very easy to detect surprising text by looking at the perplexity levels by feeding it to a very small model. If there’s a problem, OCR/‘screen read’ it instead…

    • cley_faye@lemmy.world
      link
      fedilink
      English
      arrow-up
      15
      ·
      23 days ago

      screen reader, SEO, indexation, in page search, etc.

      Basically, it breaks everything except people… unless they block/substitute fonts for accessibility reasons, in which case fuck people too.

      This is a terrible idea, and it won’t even achieve it’s original “purpose” as it is trivially detectable. Only negatives in this.

    • gex@lemmy.world
      link
      fedilink
      English
      arrow-up
      11
      ·
      23 days ago

      Yes, the decoy text is marked aria-hidden, so it won’t be read out loud. The real text is sent to the browser encrypted, and the decryption process takes ~20 seconds, roughly the same as running ocr.

      • T156@lemmy.world
        link
        fedilink
        English
        arrow-up
        7
        ·
        23 days ago

        Presumably the AI scraper would also have OCR, and would sidestep things like this?

        • sudo@programming.dev
          link
          fedilink
          English
          arrow-up
          10
          ·
          23 days ago

          A scraper has many more ways around something like sheildfont than just OCR. The question will be if it was actually programmed to check for such measures.

    • ren@reddthat.com
      link
      fedilink
      English
      arrow-up
      4
      arrow-down
      31
      ·
      23 days ago

      I bet you could even get some foaming-at-the-mouth anti-AI activists to endorse Israel’s right to resist if Israel decides to ban all AI. Worth a thought, Bibi.

      • Rothe@piefed.social
        link
        fedilink
        English
        arrow-up
        21
        arrow-down
        1
        ·
        edit-2
        23 days ago

        What a laughable strawman from a coglover. On the contrary LLM lovers will happily give money to techbro oligarchs who directly supports Trump and Israel. They will also eagerly burn down the planet just for the sake of some sloppy code.

  • merdaverse@lemmy.zip
    link
    fedilink
    English
    arrow-up
    52
    ·
    edit-2
    23 days ago

    If this gets any adoption, it will work for about a week, after which scrapers will just detect the font, and do a reverse lookup of its mapping table.

    Ironic that the repo of the font is also AI slop. If the author had asked any competent person how viable the solution is, instead of a sycophantic AI, they would have just gotten a laugh instead.

    • merdaverse@lemmy.zip
      link
      fedilink
      English
      arrow-up
      7
      ·
      23 days ago

      Thinking more about it, even if the mapping would be generated dynamically (let’s say you could generate them secretly on your server), the scraper could just parse the font file and reverse lookup the words.

    • boonhet@sopuli.xyz
      link
      fedilink
      English
      arrow-up
      5
      ·
      23 days ago

      And if you ask an AI to be critical, it’ll also tell you why this is useless lol.

      • ki4jgt@feddit.org
        link
        fedilink
        English
        arrow-up
        1
        ·
        23 days ago

        Nah. Gregg Shorthand Anniversary Edition is practically indecipherable to AI.

        It uses human intuition heavily.

  • cley_faye@lemmy.world
    link
    fedilink
    English
    arrow-up
    50
    arrow-down
    2
    ·
    23 days ago

    A terrible idea that will hinder everyone and not serve it’s original purpose in a flash.

    • anything that parse the page is broken, this includes screen reader, but also indexing, searching, and people that replace fonts locally for accessibility or other reasons
    • the “solution” for accessibility is pure trash
    • it can be trivially detected and reversed. I suspect LLM would be incredibly better at adapting to this than anything done manually too

    It’s basically a kid playing with “encrypshun” client-side, giving both the cipher and the key to the client and hoping it’ll work. Or, as other put it, DRM that don’t work for any of its original purpose, but create an additional layer of complexity and missing features, a common trend in modern projects.

  • Pika@sh.itjust.works
    link
    fedilink
    English
    arrow-up
    27
    arrow-down
    1
    ·
    23 days ago

    I want to follow that project. The only thing that I don’t like is how heavily reliant the person who runs the github is on AI-gen

    Normally, I don’t really care about it because I know that it’s something that is just a part of the industry now, but they’re using it even on their responses to people on the issue requests, and it’s to the point where it’s hurting my head trying to read it due to how drawn out and detailed it ends up being.

    It’s really hard to follow along a project where something as simple as someone opening an issue about how it doesn’t work with screen readers turns into a multi paragraph essay about the project and possibilities on how it works.

    • eyesaremosaics@lemmy.zip
      link
      fedilink
      English
      arrow-up
      1
      ·
      20 days ago

      When is Anubis useful (eg vs CloudFlare)? The GitHub page doesn’t say much-

      In most cases, you should not need this and can probably get by using Cloudflare to protect a given origin. However, for circumstances where you can’t or won’t use Cloudflare, Anubis is there for you.

  • AbouBenAdhem@lemmy.world
    link
    fedilink
    English
    arrow-up
    2
    arrow-down
    2
    ·
    edit-2
    23 days ago

    If the underlying text says one thing but the font makes it appear to say something else, which version does the author own the copyright to?

  • ki4jgt@feddit.org
    link
    fedilink
    English
    arrow-up
    2
    arrow-down
    6
    ·
    23 days ago

    Learn shorthand. AI can’t read it. Gregg Shorthand Anniversary Edition is practically indecipherable to AI. And it allows you to write at up to 300 wpm.

      • ki4jgt@feddit.org
        link
        fedilink
        English
        arrow-up
        2
        ·
        23 days ago

        The problem with that is that shorthand relies a lot on abbreviations: B = Be or By. But, at the end of a word, it can be ble, like tab. J can be J, or -age. A uses a single dot. So does -ing.

        Because of that, I’m curious whether it would catch on or not.