home.social

#aivshuman — Public Fediverse posts

Live and recent posts from across the Fediverse tagged #aivshuman, aggregated by home.social.

  1. @bignose @essjax I can confirm this.

    For starters, if you just toss a picture at an LLM, especially without a specific prompt (and alt bots don't get specific prompts), the LLM will write a generic image description. Not one specifically for the context in which the image will be posted. It won't take the expected audience into consideration either. It literally won't know the context or the audience.

    An LLM doesn't know any rules and guidelines for good image descriptions and alt-texts, especially not those specific to the Fediverse. So it won't follow any of them. (Granted, most human Fediverse users don't know any of them either. They just wing it without even knowing that they're winging it.)

    And then there's the widespread attitude that LLMs can describe any image more precisely and accurately than any human. Whoever says this has probably only ever posted stuff like real-life cat photos. But what may work for cat photos (LLMs can even fail at these) certainly will not work for extremely obscure niche content.

    What'd be extremely obscure niche content? How about renderings from 3-D virtual worlds that are so obscure that maybe one in 200,000 Fediverse users has ever even heard of the underlying technology? Which is what I'd be posting all the time if it wasn't so tedious?

    I've actually pitted an LLM that specialises in image descriptions against my own 100% human writing. Twice. In both cases, I've described an image myself, and then I've fed the same image to the LLM and prompted it to describe it for me. Then I've compared the LLM description to both my own description and the actual image. To anyone who believes the LLM wrote circles around me: No, it didn't. It failed miserably at even getting close to me.

    The LLM couldn't even tell for certain whether the image was from a video game or a virtual world. I could tell what kind of virtual world it was. In fact, I could pinpoint exactly in which location on which sim in which grid the image was taken. Not from looking at the image, but from having taken it myself. Technically, it is possible to identify at least some virtual locations by their visuals, but that'd require some super-obscure knowledge.

    The LLM left stuff out, regardless of whether it was important. In one image, there's a very prominent object that only exists in this kind of virtual worlds. The LLM ignored it completely. I described its visuals with 1,000 characters and then explained what it is, what it does, how it works and what it is for in another 4,000. In this case, there was no "important in the context" because the image itself was the very context. Besides, someone somewhere out there in the Fediverse might need this information.

    Not to mention the stuff that the LLM hallucinated. In the other description, it spoke of "a large, curved roof that appears to be made of a reflective material, possibly glass or polished metal" when describing only the end piece of said roof that has a metal-like texture with gloss effects painted on, and that isn't reflective in any way. It spoke of the four columns that support the end piece being "evenly spaced" although they're not. It spoke of "a well-maintained lawn and a few trees" when describing rather bumpy ground with a fairly basic, remotely grass-like texture painted on. It spoke of the entrance to the building not being visible in the image although a whole of three entrances are with another one being visible through the glass façade. It spoke of a sky that's "clear with a few scattered clouds" while about one third of the sky is actually covered by one large cloud.

    I got all these things right in my 100% hand-written descriptions. I managed to pinpoint exactly where the image was taken and explain the location in-depth.

    #Long #LongPost #CWLong #CWLongPost #AltText #AltTextMeta #CWAltTextMeta #ImageDescription #ImageDescriptions #ImageDescriptionMeta #CWImageDescriptionMeta #AI #LLM #LLMs #AIVsHuman #HumanVsAI
  2. @bignose @essjax I can confirm this.

    For starters, if you just toss a picture at an LLM, especially without a specific prompt (and alt bots don't get specific prompts), the LLM will write a generic image description. Not one specifically for the context in which the image will be posted. It won't take the expected audience into consideration either. It literally won't know the context or the audience.

    An LLM doesn't know any rules and guidelines for good image descriptions and alt-texts, especially not those specific to the Fediverse. So it won't follow any of them. (Granted, most human Fediverse users don't know any of them either. They just wing it without even knowing that they're winging it.)

    And then there's the widespread attitude that LLMs can describe any image more precisely and accurately than any human. Whoever says this has probably only ever posted stuff like real-life cat photos. But what may work for cat photos (LLMs can even fail at these) certainly will not work for extremely obscure niche content.

    What'd be extremely obscure niche content? How about renderings from 3-D virtual worlds that are so obscure that maybe one in 200,000 Fediverse users has ever even heard of the underlying technology? Which is what I'd be posting all the time if it wasn't so tedious?

    I've actually pitted an LLM that specialises in image descriptions against my own 100% human writing. Twice. In both cases, I've described an image myself, and then I've fed the same image to the LLM and prompted it to describe it for me. Then I've compared the LLM description to both my own description and the actual image. To anyone who believes the LLM wrote circles around me: No, it didn't. It failed miserably at even getting close to me.

    The LLM couldn't even tell for certain whether the image was from a video game or a virtual world. I could tell what kind of virtual world it was. In fact, I could pinpoint exactly in which location on which sim in which grid the image was taken. Not from looking at the image, but from having taken it myself. Technically, it is possible to identify at least some virtual locations by their visuals, but that'd require some super-obscure knowledge.

    The LLM left stuff out, regardless of whether it was important. In one image, there's a very prominent object that only exists in this kind of virtual worlds. The LLM ignored it completely. I described its visuals with 1,000 characters and then explained what it is, what it does, how it works and what it is for in another 4,000. In this case, there was no "important in the context" because the image itself was the very context. Besides, someone somewhere out there in the Fediverse might need this information.

    Not to mention the stuff that the LLM hallucinated. In the other description, it spoke of "a large, curved roof that appears to be made of a reflective material, possibly glass or polished metal" when describing only the end piece of said roof that has a metal-like texture with gloss effects painted on, and that isn't reflective in any way. It spoke of the four columns that support the end piece being "evenly spaced" although they're not. It spoke of "a well-maintained lawn and a few trees" when describing rather bumpy ground with a fairly basic, remotely grass-like texture painted on. It spoke of the entrance to the building not being visible in the image although a whole of three entrances are with another one being visible through the glass façade. It spoke of a sky that's "clear with a few scattered clouds" while about one third of the sky is actually covered by one large cloud.

    I got all these things right in my 100% hand-written descriptions. I managed to pinpoint exactly where the image was taken and explain the location in-depth.

    #Long #LongPost #CWLong #CWLongPost #AltText #AltTextMeta #CWAltTextMeta #ImageDescription #ImageDescriptions #ImageDescriptionMeta #CWImageDescriptionMeta #AI #LLM #LLMs #AIVsHuman #HumanVsAI
  3. @bignose @essjax I can confirm this.

    For starters, if you just toss a picture at an LLM, especially without a specific prompt (and alt bots don't get specific prompts), the LLM will write a generic image description. Not one specifically for the context in which the image will be posted. It won't take the expected audience into consideration either. It literally won't know the context or the audience.

    An LLM doesn't know any rules and guidelines for good image descriptions and alt-texts, especially not those specific to the Fediverse. So it won't follow any of them. (Granted, most human Fediverse users don't know any of them either. They just wing it without even knowing that they're winging it.)

    And then there's the widespread attitude that LLMs can describe any image more precisely and accurately than any human. Whoever says this has probably only ever posted stuff like real-life cat photos. But what may work for cat photos (LLMs can even fail at these) certainly will not work for extremely obscure niche content.

    What'd be extremely obscure niche content? How about renderings from 3-D virtual worlds that are so obscure that maybe one in 200,000 Fediverse users has ever even heard of the underlying technology? Which is what I'd be posting all the time if it wasn't so tedious?

    I've actually pitted an LLM that specialises in image descriptions against my own 100% human writing. Twice. In both cases, I've described an image myself, and then I've fed the same image to the LLM and prompted it to describe it for me. Then I've compared the LLM description to both my own description and the actual image. To anyone who believes the LLM wrote circles around me: No, it didn't. It failed miserably at even getting close to me.

    The LLM couldn't even tell for certain whether the image was from a video game or a virtual world. I could tell what kind of virtual world it was. In fact, I could pinpoint exactly in which location on which sim in which grid the image was taken. Not from looking at the image, but from having taken it myself. Technically, it is possible to identify at least some virtual locations by their visuals, but that'd require some super-obscure knowledge.

    The LLM left stuff out, regardless of whether it was important. In one image, there's a very prominent object that only exists in this kind of virtual worlds. The LLM ignored it completely. I described its visuals with 1,000 characters and then explained what it is, what it does, how it works and what it is for in another 4,000. In this case, there was no "important in the context" because the image itself was the very context. Besides, someone somewhere out there in the Fediverse might need this information.

    Not to mention the stuff that the LLM hallucinated. In the other description, it spoke of "a large, curved roof that appears to be made of a reflective material, possibly glass or polished metal" when describing only the end piece of said roof that has a metal-like texture with gloss effects painted on, and that isn't reflective in any way. It spoke of the four columns that support the end piece being "evenly spaced" although they're not. It spoke of "a well-maintained lawn and a few trees" when describing rather bumpy ground with a fairly basic, remotely grass-like texture painted on. It spoke of the entrance to the building not being visible in the image although a whole of three entrances are with another one being visible through the glass façade. It spoke of a sky that's "clear with a few scattered clouds" while about one third of the sky is actually covered by one large cloud.

    I got all these things right in my 100% hand-written descriptions. I managed to pinpoint exactly where the image was taken and explain the location in-depth.

    #Long #LongPost #CWLong #CWLongPost #AltText #AltTextMeta #CWAltTextMeta #ImageDescription #ImageDescriptions #ImageDescriptionMeta #CWImageDescriptionMeta #AI #LLM #LLMs #AIVsHuman #HumanVsAI
  4. @bignose @essjax I can confirm this.

    For starters, if you just toss a picture at an LLM, especially without a specific prompt (and alt bots don't get specific prompts), the LLM will write a generic image description. Not one specifically for the context in which the image will be posted. It won't take the expected audience into consideration either. It literally won't know the context or the audience.

    An LLM doesn't know any rules and guidelines for good image descriptions and alt-texts, especially not those specific to the Fediverse. So it won't follow any of them. (Granted, most human Fediverse users don't know any of them either. They just wing it without even knowing that they're winging it.)

    And then there's the widespread attitude that LLMs can describe any image more precisely and accurately than any human. Whoever says this has probably only ever posted stuff like real-life cat photos. But what may work for cat photos (LLMs can even fail at these) certainly will not work for extremely obscure niche content.

    What'd be extremely obscure niche content? How about renderings from 3-D virtual worlds that are so obscure that maybe one in 200,000 Fediverse users has ever even heard of the underlying technology? Which is what I'd be posting all the time if it wasn't so tedious?

    I've actually pitted an LLM that specialises in image descriptions against my own 100% human writing. Twice. In both cases, I've described an image myself, and then I've fed the same image to the LLM and prompted it to describe it for me. Then I've compared the LLM description to both my own description and the actual image. To anyone who believes the LLM wrote circles around me: No, it didn't. It failed miserably at even getting close to me.

    The LLM couldn't even tell for certain whether the image was from a video game or a virtual world. I could tell what kind of virtual world it was. In fact, I could pinpoint exactly in which location on which sim in which grid the image was taken. Not from looking at the image, but from having taken it myself. Technically, it is possible to identify at least some virtual locations by their visuals, but that'd require some super-obscure knowledge.

    The LLM left stuff out, regardless of whether it was important. In one image, there's a very prominent object that only exists in this kind of virtual worlds. The LLM ignored it completely. I described its visuals with 1,000 characters and then explained what it is, what it does, how it works and what it is for in another 4,000. In this case, there was no "important in the context" because the image itself was the very context. Besides, someone somewhere out there in the Fediverse might need this information.

    Not to mention the stuff that the LLM hallucinated. In the other description, it spoke of "a large, curved roof that appears to be made of a reflective material, possibly glass or polished metal" when describing only the end piece of said roof that has a metal-like texture with gloss effects painted on, and that isn't reflective in any way. It spoke of the four columns that support the end piece being "evenly spaced" although they're not. It spoke of "a well-maintained lawn and a few trees" when describing rather bumpy ground with a fairly basic, remotely grass-like texture painted on. It spoke of the entrance to the building not being visible in the image although a whole of three entrances are with another one being visible through the glass façade. It spoke of a sky that's "clear with a few scattered clouds" while about one third of the sky is actually covered by one large cloud.

    I got all these things right in my 100% hand-written descriptions. I managed to pinpoint exactly where the image was taken and explain the location in-depth.

    #Long #LongPost #CWLong #CWLongPost #AltText #AltTextMeta #CWAltTextMeta #ImageDescription #ImageDescriptions #ImageDescriptionMeta #CWImageDescriptionMeta #AI #LLM #LLMs #AIVsHuman #HumanVsAI