• Kuvwert@lemm.ee
      link
      fedilink
      English
      arrow-up
      4
      ·
      1 day ago

      Non thinking prediction models can’t count the r’s in strawberry due to the nature of tokenization.

      However openai o1 and deep seek r1 can both reliably do it correctly

    • blakenong@lemmings.world
      link
      fedilink
      English
      arrow-up
      9
      arrow-down
      1
      ·
      edit-2
      1 day ago

      No. It literally cannot count the number of R letters in strawberry. It says 2, there are 3. ChatGPT had this problem, but it seems it is fixed. However if you say “are you sure?” It says 2 again.

      Ask ChatGPT to make an image of a cat without a tail. Impossible. Odd, I know, but one of those weird AI issues

      • SoftestSapphic@lemmy.world
        link
        fedilink
        English
        arrow-up
        6
        arrow-down
        3
        ·
        1 day ago

        Because there aren’t enough pictures of tail-less cats out there to train on.

        It’s literally impossible for it to give you a cat with no tail because it can’t find enough to copy and ends up regurgitating cats with tails.

        Same for a glass of water spilling over, it can’t show you an overfilled glass of water because there aren’t enough pictures available for it to copy.

        This is why telling a chatbot to generate a picture for you will never be a real replacement for an artist who can draw what you ask them to.

        • JustARaccoon@lemmy.world
          link
          fedilink
          English
          arrow-up
          4
          arrow-down
          1
          ·
          1 day ago

          Not really it’s supposed to understand what a tail is, what a cat is, and which part of the cat is the tail. That’s how the “brain” behind AI works

          • SoftestSapphic@lemmy.world
            link
            fedilink
            English
            arrow-up
            2
            arrow-down
            3
            ·
            edit-2
            1 day ago

            It searches the internet for cats without tails and then generates an image from a summary of what it finds, which contains more cats with tails than without.

            That’s how this Machine Learning progam works

            • Kogasa@programming.dev
              link
              fedilink
              English
              arrow-up
              2
              arrow-down
              1
              ·
              12 hours ago

              It doesn’t search the internet for cats, it is pre-trained on a large set of labelled images and learns how to predict images from labels. The fact that there are lots of cats (most of which have tails) and not many examples of things “with no tail” is pretty much why it doesn’t work, though.

                • Kogasa@programming.dev
                  link
                  fedilink
                  English
                  arrow-up
                  2
                  arrow-down
                  1
                  ·
                  10 hours ago

                  It’s not the “where” specifically I’m correcting, it’s the “when.” The model is trained, then the query is run against the trained model. The query doesn’t involve any kind of internet search.

                  • SoftestSapphic@lemmy.world
                    link
                    fedilink
                    English
                    arrow-up
                    1
                    arrow-down
                    2
                    ·
                    9 hours ago

                    And I care about “how” it works and “what” data it uses because I don’t have to walk on eggshells to preserve the sanctity of an autocomplete software

                    You need to curb your pathetic ego and really think hard about how feeding the open internet to an ML program with a LLM slapped onto it is actually any more useful than the sum of its parts.

            • FatCrab@lemmy.one
              link
              fedilink
              English
              arrow-up
              2
              arrow-down
              1
              ·
              16 hours ago

              That isn’t at all how something like a diffusion based model works actually.

              • SoftestSapphic@lemmy.world
                link
                fedilink
                English
                arrow-up
                1
                arrow-down
                1
                ·
                16 hours ago

                So what training data does it use?

                They found data to train it that isn’t just the open internet?

                • FatCrab@lemmy.one
                  link
                  fedilink
                  English
                  arrow-up
                  2
                  arrow-down
                  1
                  ·
                  14 hours ago

                  Regardless of training data, it isn’t matching to anything it’s found and squigglying shit up or whatever was implied. Diffusion models are trained to iteratively convert noise into an image based on text and the current iteration’s features. This is why they take multiple runs and also they do that thing where the image generation sort of transforms over multiple steps from a decreasingly undifferentiated soup of shape and color. My point was that they aren’t doing some search across the web, either externally or via internal storage of scraped training data, to “match” your prompt to something. They are iterating from a start of static noise through multiple passes to a “finished” image, where each pass’s transformation of the image components is a complex and dynamic probabilistic function built from, but not directly mapping to in any way we’d consider it, the training data.

                  • SoftestSapphic@lemmy.world
                    link
                    fedilink
                    English
                    arrow-up
                    1
                    arrow-down
                    1
                    ·
                    edit-2
                    13 hours ago

                    Oh ok so training data doesn’t matter?

                    It can generate any requested image without ever being trained?

                    Or does data not matter when it makes your agument invalid?

                    Tell me how you moving the bar proves that AI is more intelligent than the sum of its parts?

        • vrighter@discuss.tchncs.de
          link
          fedilink
          English
          arrow-up
          2
          arrow-down
          2
          ·
          22 hours ago

          so… with all the supposed reasoning stuff they can do, and supposed “extrapolation of knowledge” they cannot figure out that a tail is part of a cat, and which part it is.

          • Kuvwert@lemm.ee
            link
            fedilink
            English
            arrow-up
            2
            ·
            10 hours ago

            The “reasoning” models and the image generation models are not the same technology and shouldn’t be compared against the same baseline.

            • vrighter@discuss.tchncs.de
              link
              fedilink
              English
              arrow-up
              1
              arrow-down
              2
              ·
              12 hours ago

              I’m not seeing any reasoning, that was the point of my comment. That’s why I said “supposed”