• Blackmist@feddit.uk
    link
    fedilink
    English
    arrow-up
    7
    ·
    9 days ago

    Well I just tested it with something it should know if it was doing that (some API documentation we don’t publicly release), and it just made up some complete nonsense.

    So… Maybe not?

    • Nurse_Robot@lemmy.world
      link
      fedilink
      English
      arrow-up
      6
      ·
      9 days ago

      If it does train, that doesn’t mean it will use that data immediately or consistently reference one data point. I don’t think your test proves or disproves anything

      • Blackmist@feddit.uk
        link
        fedilink
        English
        arrow-up
        4
        ·
        9 days ago

        Probably. The document is specific to my software, and has been there a long time. Long enough that it should be in the training data, although getting out to pull that particular bit out could be a pain.

        Maybe they only trained on documents with certain privacy options set, or use some heuristics to determine if they should use it or not. Maybe the person in the article had his data hacked and shared on a fan site.

        Hard to tell, and it’s not like Google are going to tell us either way.