home.social

#imagedescriptionmeta — Public Fediverse posts

Live and recent posts from across the Fediverse tagged #imagedescriptionmeta, aggregated by home.social.

  1. CW: Account getting away with way subpar alt-texts only because it's too niche with too few followers; CW: long (over 1,700 characters), Fediverse meta, alt-text meta, image description meta, character limit meta
    I've just discovered a certain Mastodon account that seems to automatically post images of Second Life avatars. It can be lucky to only have eight followers, one of them being a search engine, the others being Second Life users.

    If it had more follower, it would certainly already have met the wrath of the wider Mastodon community, especially the Mastodon HOA, for breaking Mastodon's unwritten rules.

    And I'm not even talking about nigh-nudity in some of the images with no warning, no image flagging, no hashtag. I'm talking about the alt-texts that are vastly below Mastodon's requirements and quality standards for alt-texts.

    Granted, all things considered, the requirements for good image descriptions by Mastodon's standards have to be extremely hard and tedious to meet, at least by my personal estimations. In fact, they have to be impossible to meet on Mastodon itself due to its tiny character limit, and if this account was somewhere where it could meet these requirements, it'd probably be blocked by loads of Mastodon accounts for its excessively long posts due to the long image descriptions.

    But if more Mastodon users knew it right now, it'd probably be blocked by many more Mastodon accounts than follow it.

    If you post from virtual worlds into the Fediverse, you simply cannot win in the long run. You'll lose either way.

    #SecondLife #Metaverse #VirtualWorld #VirtualWorlds #Long #LongPost #CWLong #CWLongPost #FediMeta #FediverseMeta #CWFediMeta #CWFediverseMeta #AltText #AltTextMeta #CWAltTextMeta #ImageDescription #ImageDescriptions #ImageDescriptionMeta #CWImageDescriptionMeta #CharacterLimit #CharacterLimits #CharacterLimitMeta #CWCharacterLimitMeta #Mastodon #MastodonHOA #MastodonCulture
  2. @Robert Kingett I've read somewhere that some blind people prefer LLM-generated image descriptions to human-written image descriptions because they're more entertaining. LLM-generated image descriptions often add some whimsy whereas human-written image descriptions just dryly rattle down what's in the image. Basically, through the screen reader, humans sound more like machines than actual machines.

    They usually are aware that LLMs tend to be hallucinating and essentially telling them non-sense. But I've read from one blind user that they don't care whether or not the LLM-generated description is accurate as long as the accuracy isn't a matter of life and death.

    In a certain way, it is understandable. The only people who'd criticise an image description for being inaccurate are fully sighted and therefore capable of comparing the image with its description.

    Blind people won't notice unless the image description describes something so outlandish to them that they have the impression of an utterly surrealist image where there shouldn't be a surrealist image.

    On the other hand, there are also blind people who demand image descriptions be accurate. They simply don't want to be told non-sense, especially not without knowing that they're being told non-sense.

    But even on the sighted side, there's the "human versus AI" debate.

    Some sighted people are fully convinced that LLMs can describe absolutely every image perfectly in absolutely every situation, no matter how obscure the contents of the image are. They're fully convinced that a generic LLM like ChatGPT can write circles around even human experts at any given time.

    Some are simply AI fanbois or fangurls. Others say so in order to convince themselves that what they're doing is the best way: They use image-describing LLMs or even general-purpose LLMs as fire-and-forget tools. They have image descriptions generated, they copy-paste these image descriptions into the alt-texts, they send their posts, and they never take a look at these image descriptions at any point in the process. It's more convenient this way.

    And then they wonder why they're under attack from sighted alt-text activists who call them out for their painfully inaccurate and blatantly obvious AI slop.

    LLM proponents, both sighted and non-sighted, have in common that they never compare the image and the description. Non-sighted LLM proponents simply can't see the image. Sighted LLM proponents put so much faith into LLMs that they can't be bothered to read the description and cross-check it with the image.

    #Long #LongPost #CWLong #CWLongPost #AltText #AltTextMeta #CWAltTextMeta #ImageDescription #ImageDescriptions #ImageDescriptionMeta #CWImageDescriptionMeta #AI #LLM #LLMs #AIVsHuman #HumanVsAI
  3. @hosh The alt-text police of the Mastodon Home Owners' Association (Mastodon HOA) have a tendency to be overzealous. And they don't talk to each other. They all act for themselves as lone wolves with exactly no coordination amongst each other whatsoever. You never know what kinds of rules they whip up for themselves.

    Chances are that they only let image descriptions count that come directly with the image. If they acknowledge an image description outside the alt-text, it must be in the post itself. Not as an external link, but the description text itself.

    Besides, at least some in the Mastodon HOA have problems with external links. And I don't just mean that they don't trust embedded links whose URL they can't see in plain sight, the kind that Hubzilla can create and Mastodon can't (not that Hubzilla couldn't fake a plain-sight link by embedding a different URL than the visible).

    I don't mean either that probably a majority of Mastodon users don't even recognise embedded links without a visible URL as such because they don't know that such a thing can exist in the Fediverse, because Mastodon can't make them.

    No, what I mean is the notion that external links for explanations are inherently bad from an accessibility point of view. "Mastodon" (as in how Mastodon users experience the Fediverse, i.e. the Mastodon Web UI or any of the popular mobile phone apps) is sufficiently accessible. But the Web outside of "Mastodon" (same definition again) may not be accessible enough.

    A few years ago, I've literally read a Mastodon toot in which someone said that explanations must not be linked to. Linked websites have a risk of not being accessible. Explanations must always be directly in the same post. Apparently, they thought that everything and anything can be explained and broken down until everyone understands it within 500 characters.

    This is also why Mastodon users tend to explain their images in the alt-text. It's only there where they have at least halfway enough characters for an explanation, 1,500 per image as opposed to usually only 500 in the post text. (On Mastodon, much unlike Hubzilla, the alt-text is a separate database field that exists separately for each of the up to four images per message.)

    That is, explanations must never go into the alt-text because there are people who cannot open alt-texts to read them. But nobody on Mastodon knows that.

    It should be obvious that what counts for explanations counts for visual descriptions just as well.

    And in fact, regarding Hubzilla articles, they're actually right. I've once pointed an actually blind screen reader user to an article on my Hubzilla channel. She said she couldn't even navigate the Web interface. She literally couldn't get to the text body of the article to have it read out by her screen reader.

    Hubzilla's Web interface, no matter which app is opened, is not accessible. It does not work with screen readers. It's largely still stuck in 2012 when nobdy made any ruckus about the accessibility of hobbyist Web projects.

    The only reason why at least some blind or visually-impaired users can read our Hubzilla posts and comments and DMs is because they're all on Mastodon, and they read our content either on Mastodon's Web UI or a Mastodon app that supports screen readers. But they do not read our content at the source. Because they can't.

    I actually took into consideration linking to my long image descriptions. But my idea was not to link to a Hubzilla article, nor to a Hubzilla wiki or a Hubzilla card. No, my idea was to write a plain HTML document, upload it to my file space and link to that.

    I've dropped that idea for various reasons:
    • Generally, still, external links are frowned upon.
    • I don't know if plain HTML is accessible without a CSS. And I can't add a CSS to this HTML if the HTML document is not served to the recipient by a Web server, but by a file server.
    • I don't know how Google Chrome on Android or Safari on an iPhone will react when they access an HTML document on a file space. Will they display it as a website? Or will they download it onto the device as a file without opening it because, again, it is served to them not by a Web server, but by a file server?
    • Mobile users dislike opening websites from apps because they dislike their browser popping open. And on Mastodon, much unlike Hubzilla, almost everyone is on a phone and a dedicated app almost all the time.
    • This also means that mobile users would have the image and the description in two separate apps. The image in their Mastodon app, the description in their browser.
    • I would need much more description.
      Right now, when I have multiple images, my long descriptions consist of a preamble that contains all necessary explanations and, if applicable, visual descriptions of elements that are common to all images. The individual descriptions for each image follow.
      But if I had one image description file per image, then each image description would need the whole preamble included. I can't just add the preamble to the first description file.
      What if someone opens the third description file first? They'll only have a very incomplete description. And linking to the first description file is inconvenient. I would have to know the URL of the first file before completing and uploading the other files because I'd have to include the URL of the first file in them. And the users would have to have three documents open (the image post, the description of the image they're interested in, the description of the first image with the preamble) just to experience one image. Spread across two phone apps.

    And that's why I can't put my additional long description in an external document.

    #Long #LongPost #CWLong #CWLongPost #FediMeta #FediverseMeta #CWFediMeta #CWFediverseMeta #AltText #AltTextMeta #CWAltTextMeta #ImageDescription #ImageDescriptions #ImageDescriptionMeta #CWImageDescriptionMeta #AltTextPolice #MastodonHOA #CharacterLimit #CharacterLimits #CharacterLimitMeta #CWCharacterLimitMeta #500Characters #MastodonCulture
  4. CW: Long posts vs insufficiently described images: I can't win either way; CW: long (over 3,500 characters), Fediverse meta, Fediverse-beyond-Mastodon meta, alt-text meta, image description meta, content warning meta, character limit meta
    On the one hand, I have to go out of my way and write two image descriptions for each one of my original images. One is short and goes into the alt-text, and I'm going to limit all my future alt-text to a maximum of 512 characters (otherwise users on Misskey, Sharkey etc. will believe I haven't written any alt-text because they won't receive any due to a bug).

    The other one is enormous degrees of magnitudes longer than anything most Fediverse users have ever read in the Fediverse. It also contains all explanations necessary to understand the image and its description, and if there's text anywhere within the borders of the image, readable or not, it contains verbatim transcripts of said text.

    The nature of my original images requires such long descriptions. Besides, the only way to really be safe from the alt-text police of the Mastodon HOA is to overcomply with whatever minimum standards for good image descriptions anyone of them may have.

    On the other hand, the self-same Mastodon HOA is likely to sanction me for the self-same posts. The reason: The posts are way too long. They exceed the limit of 500 characters that's so deeply engrained into Mastodon's culture that many Mastodonians are eager to defend it. Even if I hide them behind a summary with a content warning about the post being long. If I were to appease these Mastodonians, I'd have to underdescribe my images, and I wouldn't be able to explain them at all.

    Speaking of underdescribing, I think at least some members of the alt-text police actually don't let image descriptions in the post count. What counts is only the image description in the alt-text. It must be accurate, it must be sufficiently detailed, and it must contain all the text transcripts. In fact, I wouldn't wonder if they demanded sufficient explanations in the alt-text, not knowing that explanations in alt-text are actually a big no-no.

    Even if all requirements of a good alt-text by alt-text police standards are met or even exceeded by the image description in the post, chances are the alt-text police will still sanction me if the alt-text doesn't meet these criteria.

    When it comes to my original images, even squeezing all that into the 1,500-character limit for alt-texts imposed by Mastodon is pretty much impossible. Squeezing it into the 512-character limit for alt-text imposed by Misskey and its forks is even more impossible.

    The only winning move is to not play at all. Curiously, some people are even upset about me rarely posting any images. Although they don't follow me. Although the channel that I use for original images (@Jupiter Rowland's (streams) outlet) has next to no reach, so even if I were to post images again, practically nobody would notice. Although it doesn't even seem that there's much interest in that kind of images in the first place.

    But apparently, according to some, posting images with only rudimentary alt-text whipped up in a minute, no long description and no explanations is always so much better than not posting images because it takes so much time and effort to describe them.

    #Long #LongPost #CWLong #CWLongPost #FediMeta #FediverseMeta #CWFediMeta #CWFediverseMeta #AltText #AltTextMeta #CWAltTextMeta #ImageDescription #ImageDescriptions #ImageDescriptionMeta #CWImageDescriptionMeta #AltTextPolice #MastodonHOA #CW #CWs #CWMeta #ContentWarning #ContentWarnings #ContentWarningMeta #CWContentWarningMeta #CharacterLimit #CharacterLimits #CharacterLimitMeta #CWCharacterLimitMeta #500Characters #MastodonCulture
  5. @Cassandrich @Sobri | Zoe (she/her) @Scott Jenson @Phil Dennis-Jordan Also, an image doesn't always need the exact same alt-text whenever it's posted somewhere.

    The alt-text must adapt to the context. It must be different according to the context in which an image is posted. Also, it must adapt to the place where it's posted. The same image, even within a very similar context, must have a different alt-text in the Fediverse than on commercial social media or a static website. Lastly, and this ties in with the Fediverse requiring different alt-texts, the audience must be taken into consideration.

    Alt-text in metadata can't do either of this. An LLM can't do either of this either unless it's explicitly prompted to do so, and even that is questionable.

    Many Mastodon users dream of only pressing a button or not even that, and some AI automagically generates a perfect alt-text for their image. Perfectly accurate with exactly the details required for the context and the intended audience as well as the expected audience, all while following every last image description and alt-text rule out there to a tee.

    It's perfectly understandable. Mastodon had begun to feel like child's play when they were suddenly pressured into describing each and every image they post. Worse yet, it seems like over 90% of all Mastodon users do everything on a phone with no access to a hardware keyboard whatsoever. So they have to fumble their alt-texts into a screen keyboard while not even being able to see the image they're describing.

    I'm neither on Mastodon nor on a phone. I've got the luxury of having a desktop computer with a hardware keyboard and being able to bllind-type. So I don't have a problem with writing my image descriptions myself with no help from an AI.

    In fact, my own original images are all about an extreme niche topic. It's so obscure that no AI will ever be able to describe such images, much less explain them at my level of accuracy and detail. (Explanations go into the post text, by the way, and not into the alt-text, but I always have an additional image description in the post text for my original images anyway.)

    I simply know things that no AI will ever know, not ChatGPT and not Claude either, at least not at the point in time when they need that knowledge. And I can see things that will always remain invisible for AIs.

    You can develop better models all you want. But they'll never be able to do all that.

    #Long #LongPost #CWLong #CWLongPost #FediMeta #FediverseMeta #CWFediMeta #CWFediverseMeta #AltText #AltTextMeta #CWAltTextMeta #ImageDescription #ImageDescriptions #ImageDescriptionMeta #CWImageDescriptionMeta #AI #AIVsHuman #HumanVsAI
  6. @C. I have two major issues with the Mastodon HOA.

    One, they try hard to force "Mastodon standards", Mastodon culture and Mastodon's unwritten rules upon the whole Fediverse. Including places that not only aren't Mastodon, but that are very much not Mastodon. Simply because they can't see where a message is from. In fact, many of them are still fully convinced that the Fediverse is only Mastodon.

    And so you have members of the Mastodon HOA yelling at someone who is allegedly "doing Mastodon wrong", but that someone is actually on Friendica and has been since as early as 2011. As in about five years longer than Mastodon has even existed. And seriously, the only places in the Fediverse that are even more different and farther away from Mastodon than Friendica (without specialising in something that Mastodon absolutely can't do) are Friendica's own descendants: Hubzilla, (streams), Forte.

    The Mastodon HOA probably don't know that Friendica exists. They definitely don't know that either of the other three exists. They definitely don't know that any of the four is significantly different from Mastodon in any way. And frankly, they don't care a bit. If it appears on any Mastodon timeline, it's Mastodon to them, and it has to adapt to Mastodon's culture and follow Mastodon's rules.

    Two, they don't coordinate anything among each other. They're just a bunch of lone wolves. Everyone has got their own standards, but everyone thinks their personal standards are the one and only Mastodon/Fediverse gold standards, and everyone enforces their own standards. And, of course, everyone thinks their standards can and must apply always, including in the most obscure edge-cases.

    For example, they've got standards for describing real-life photos on Mastodon with a character limit of 500. And they try to enforce these standards always and everywhere. However, these standards don't necessarily work perfectly when I post a rendering from a super-obscure 3-D virtual world on (streams) with a character limit of over 24 million where I've got loads of room to write an additional long image description and put it into the post text.

    The Mastodon HOA, or at least some of their members, appear to be constantly raising their minimum quality requirements for image descriptions. They must be absolutely accurate, and they must be sufficiently detailed that nobody will ever have to ask for a detail description. Oh, and they must explain whatever the audience may not know about the image or the description. (At this point, it's fair to mention that explanations must never go into the alt-text.)

    Sure, I can do that. I have done so in the past. But I can't do that within Mastodon's alt-text character limit of 1,500 (Mastodon truncates longer alt-texts from outside). I can do that even less within Misskey's alt-text character limit of only 512 (Misskey and the Forkeys should truncate longer alt-texts, but due to a bug, they delete them entirely instead, giving the impression that you haven't written an alt-text at all). I can only do that in the additional long description in the post text.

    If the Mastodon HOA demand I transcribe literally any and all text within the borders of an image, I can do that, too. In fact, I have done so in the past. I can transcribe bits of text verbatim which the Mastodon HOA can't even read. Which the Mastodon HOA couldn't even find in the image because they're so tiny. But there's no way that I can squeeze 20+ individual text transcripts into 1,500 characters of alt-text along with the rest of the visual description, much less into only 512 characters. The text transcripts will have to go into the long description in the post text, whether the Mastodon HOA want or not.

    This means that the post will exceed the holy limit of 500 characters by huge magnitudes. This, in turn, means that when I've satisfied one Mastodon HOA member, another one comes and sanctions me for exceeding the holy 500-character limit. That is, chances are it's actually the same Mastodon HOA member.

    In other words, if the content of an image is obscure enough and requires enough description, the only winning move when I want to post such an image is to not post it at all.

    #Long #LongPost #CWLong #CWLongPost #FediMeta #FediverseMeta #CWFediMeta #CWFediverseMeta #CharacterLimit #CharacterLimits #CharacterLimitMeta #CWCharacterLimitMeta #500Characters #AltText #AltTextMeta #CWAltTextMeta #ImageDescription #ImageDescriptions #ImageDescriptionMeta #CWImageDescriptionMeta #MastodonCulture #MastodonHOA
  7. @Woochancho @Diego Martínez (Kaeza) 🇺🇾 @🅰🅻🅸🅲🅴  (🌈🦄) Especially whenever humans have advantages over LLMs.

    When I describe my own original images, I have two advantages.

    One, I know much more about the contents of the image than any AI. That's because my original images always show something from extremely obscure 3-D virtual worlds. On top of that, I may add some extra insider knowledge or explain pop-cultural references in the long description in the post if it helps understand the image and its descriptions.

    Two, the LLM can only look at the image with its limited resolution. That's all it has. In contrast, when I describe my images, I don't just look at the images. I look at the real deal in-world with a nearly infinite resolution.

    For example, an LLM can only generate a description from a picture of a virtual building. But when I describe it, my avatar is in-world, standing right in front of the building whose picture I'm describing. I can move the avatar around, I can move the camera around, I can zoom in on anything. I can correctly identify that four-pixel blob as a strawberry cocktail wheras the LLM doesn't even notice it's there.

    I've actually done two tests using LLaVA. I've fed it two images I had described myself previously to see what happens. It was abysmal. LLaVA hallucinated, it interpreted stuff wrongly and so forth, not to mention that LLaVA's description, even after being prompted to write a detailed description, wasn't nearly as detailed as mine.

    In one image, there's an OpenSimWorld beacon placed rather prominently in the scenery. LLaVA completely ignored it. I described what it looks like in about 1,000 characters, and then I explained what it is, what OpenSimWorld is and how it works in another 4,000 characters or so.

    It's an illusion that AI will soon catch up with any of this.

    Oh, by the way: How is an AI supposed to pinpoint exactly where an image was made if the image shows a place of which multiple absolutely identical copies exist? Or if the image has a neutral background that doesn't even hint at where it was made? I can do that with no problem because I remember where I've made the image.

    #Long #LongPost #CWLong #CWLongPost #AltText #AltTextMeta #CWAltTextMeta #ImageDescription #ImageDescriptions #ImageDescriptionMeta #CWImageDescriptionMeta #AI #LLaVA #AIVsHuman #HumanVsAI
  8. @Pino Carafa Well, my problem is not the alt-text.

    I used to limit my alt-texts to 1,500 characters because Mastodon and its forks truncate longer alt-texts at the 1,500-character mark. In the future, I will limit them to 512 characters because Misskey and its forks should truncate them at that mark if they're longer, but instead, they delete them.

    But in addition to my alt-texts, I describe my original images once more (= twice altogether). The other description is what I call the "long description", and it goes directly into the post text (as opposed to the alt-text). I don't have a character limit to worry about (over 16.7 million), so I can do what's outright unimaginable from a Mastodon point of view.

    It's this long description that's causing trouble.

    That is, I wouldn't wonder if the Mastodon HOA were to sanction me for my alt-text not being detailed enough when I limit it to 512 characters. In fact, I wouldn't wonder if they were to sanction me because a 1,500-character alt-text of mine is lacking important elements (descriptions of certain details, transcripts of all text within the borders of the image etc.).

    #Long #LongPost #CWLong #CWLongPost #FediMeta #FediverseMeta #CWFediMeta #CWFediverseMeta #CharacterLimit #CharacterLimits #CharacterLimitMeta #CWCharacterLimitMeta #AltText #AltTextMeta #CWAltTextMeta #ImageDescription #ImageDescriptions #ImageDescriptionMeta #CWImageDescriptionMeta #MastodonHOA
  9. @Pino Carafa An additional advantage of this would be that I could first ask just how detailed a description they need. Like, if they really want me to spend two full days, morning to evening, to write something that'll take their screen reader three hours to read out loud.

    The problem, however, is that the virtual worlds that I frequent change a lot. Everything is built by users. A place that I've shown in an image may change mere days or hours after I've been there, so when I go back to take a closer look for a detailed description, it doesn't look like on the image anymore.

    Or that place may be gone entirely. For example, I could post some images from an in-world event, from places specifically built for this event. Then, two months later, someone asks for a more detailed description. But I can't write a more detailed description because I can't go back to these places, simply because these places were closed and shut down a few days after I had posted the images.

    Lastly, my impression of Mastodon is still that a significant number of users do not want to ask. Whatever information they may need, they expect it all to come with the post immediately. Having to ask for a detail description or for an explanation appears to be about as bad style as having to ask for a description in the first place.

    I've literally seen Mastodon toots in which people say that if they don't understand a post or an image in a post, they want an explanation to come with the post.

    I've also seen a Mastodon toot in which someone said that it isn't sufficient to just say what's in an image, but you also have to describe what it looks like. Right away. And in my case, this is actually absolutely justified.

    It's a catch-22: If I don't describe my images sufficiently, I risk being sanctioned by the Mastodon HOA for not describing my images sufficiently. But if I do, I risk being sanctioned by the Mastodon HOA for exceeding 500 characters in one post.

    Oh, and if I chop my image descriptions into tiny chunks of no more than 500 characters, it's disturbing for my own ilk, the users of Friendica, Hubzilla, (streams) and Forte, who are used to not having any character limits and everything being in one message, no matter how long it is. Besides, how many Mastodon users are willing to read a thread of 120 or more posts and find that more convenient than one post with 60,000 characters?

    #Long #LongPost #CWLong #CWLongPost #FediMeta #FediverseMeta #CWFediMeta #CWFediverseMeta #CharacterLimit #CharacterLimits #CharacterLimitMeta #CWCharacterLimitMeta #500Characters #AltText #AltTextMeta #CWAltTextMeta #ImageDescription #ImageDescriptions #ImageDescriptionMeta #CWImageDescriptionMeta #MastodonHOA
  10. @Mastodon Migration
    Basically telling other people how they should be using Mastodon is not cool unless they are violating some instance rule.

    As, by the way, is telling Fediverse users who are not on Mastodon to use whatever they use instead like Mastodon users are expected to use Mastodon.

    Please don't be a Mastodon HOA enforcer.

    Especially since the alt-text police of the Mastodon HOA have much higher alt-text and image description minimum standards than blind or visually-impaired people. And they seem to be raising their standards further and further.

    I always try my best to be way ahead of anyone's image description minimum standards, also in order to demonstrate to the Mastodon HOA that I'm not a lazy bum, and that I do try hard to describe my images properly. For my own original images, this means that I have to describe each one of them twice, with a fairly short description in the alt-text and a much longer one in the post itself.

    This, however, clashes with the Mastodon HOA, too, because they also enforce Mastodon's default 500-character limit Fediverse-wide by generously blocking everyone whom they catch exceeding it at first strike.

    CC: @🅰🅻🅸🅲🅴  (🌈🦄)

    #Long #LongPost #CWLong #CWLongPost #FediMeta #FediverseMeta #CWFediMeta #CWFediverseMeta #CharacterLimit #CharacterLimits #CharacterLimitMeta #CWCharacterLimitMeta #500Characters #AltText #AltTextMeta #CWAltTextMeta #ImageDescription #ImageDescriptions #ImageDescriptionMeta #CWImageDescriptionMeta #AltTextPolice #MastodonHOA
  11. @モスケ^^ ❄️🐈🔥🐴 No. Very clearly no.

    People keep thinking that AI solves the alt-text problem perfectly. Like, push one button, get a perfect alt-text for your image, send it without having to check it. Or, better yet, don't even push a button, the AI will take care of everything fully automatically.

    However, at best, AI-generated alt-text is better than nothing. Oftentimes, AI-generated alt-text is literally worse than nothing.

    First of all, AI does not know the context in which an image is posted. But an alt-text should always be written for a specific context because it usually depends on the context what needs to be described at all and on which level of detail.

    This means that AI tends to leave out details that may be important while describing details that literally nobody is interested in.

    AI can't take your target audience/your actual audience into consideration either. It can't write an alt-text specifically for that audience, fine-tuned for what that audience knows, what it doesn't know and what it needs and/or wants to know.

    Worse yet, AI tends to hallucinate. It tends to mention stuff in an image that simply isn't there. It tends to describe elements of an image falsely. You could post a photo of a Yorkshire terrier, and the AI may think it's a cat because it can't distinguish it from a cat in that photo.

    Seriously, AI may get even descriptions of simple images of very common things wrong. If you post images with very obscure, very niche content, AI fares even worse because it knows nothing about that very obscure, very niche content.

    If you post a screenshot from social media, AI will not necessarily know that it has to transcribe the text in the screenshot 100% verbatim. And just pushing one button or running AI on full-auto, the thing that so many smartphone users are so much craving for, will not prompt it to do so.

    If you want good, useful, accurate, sufficiently detailed image descriptions that match both the context of your posts and your audience, you will have to write them yourself.

    Trust me. I know from personal experience. I post some of the most obscure niche stuff in the Fediverse. And I've pitted an image-describing AI against my own 100% hand-written image descriptions twice already. The AI failed miserably to even come close to my descriptions in both cases.

    #Long #LongPost #CWLong #CWLongPost #AltText #AltTextMeta #CWAltTextMeta #ImageDescription #ImageDescriptions #ImageDescriptionMeta #CWImageDescriptionMeta #AI #AIVsHuman #HumanVsAI
  12. @iolaire After I have written the long description, distilled the short description from it and posted the image with both, I have asked a LLM AI for a description.

    The AI of my choice was LLaVA 1.6: https://llava.hliu.cc/

    The prompt was, "Describe the image in detail."

    LLaVA took about half a minute to generate this image description:

    The image depicts a modern architectural structure with a distinctive design. The building features a large, curved roof that appears to be made of a reflective material, possibly glass or polished metal. The roof is supported by several tall, slender columns that are evenly spaced and rise from the ground to the roof's edge. The structure has a circular emblem on the front, which includes a stylized letter 'M' and a series of concentric circles, suggesting it might be a logo or emblem of some sort.

    The building is situated on a landscaped area with a well-maintained lawn and a few trees. There is a paved walkway leading up to the entrance of the building, which is not visible in the image. The sky is clear with a few scattered clouds, indicating fair weather conditions. The overall style of the image is a digital rendering or a photograph of a 3D model, as indicated by the smooth surfaces and the absence of any visible texture or imperfections that would be present in a real-world photograph. There are no visible texts or brands that provide additional context about the building's purpose or location.


    (5/6)

    #Long #LongPost #CWLong #CWLongPost #ImageDescription #ImageDescriptions #ImageDescriptionMeta #CWImageDescriptionMeta #AI #LLaVA #AIVsHuman #HumanVsAI
  13. @iolaire Allow me to give you an example.

    This is the image I'm talking about: https://hub.netzgemeinde.eu/photos/jupiter_rowland/image/b1e7bf9c-07d8-45b6-90bb-f43e27199295 (linked instead of embedded so I don't have to go through the hassle of having to describe it right here right now).

    This is the thread in which I've posted the image before, including image descriptions, also including a comment with the AI description and an analysis of the AI description in comparison with my own descriptions: https://hub.netzgemeinde.eu/item/f8ac991d-b64b-4290-be69-28feb51ba2a7 (yes, this is part of the Fediverse; it's on the same Hubzilla channel that I'm commenting from right now).

    (2/6)

    #Long #LongPost #CWLong #CWLongPost #FediMeta #FediverseMeta #CWFediMeta #CWFediverseMeta #AltText #AltTextMeta #CWAltTextMeta #ImageDescription #ImageDescriptions #ImageDescriptionMeta #CWImageDescriptionMeta #AI #AIVsHuman #HumanVsAI
  14. @iolaire I've pitted an image-describing LLM AI against my own 100% hand-written image descriptions twice so far. I have first described an image myself, twice even, with a "short" description for the alt-text and a long, fully detailed description with text transcripts and all necessary explanations for the post text.

    However, I'm always at an unfair advantage. My images are renderings from very obscure 3-D virtual worlds. LLMs know next to nothing or actually nothing about these worlds whereas I dare say I'm an expert on them. An AI couldn't even tell whether the image is from a game or from a virtual world, much less which virtual world. I can not only exactly pinpoint where the image was taken (which place on which sim in which grid), but also explain the location and these virtual worlds in general.

    Besides, an AI would describe the image by examining the image. I describe my images by going in-world and looking at the real deal instead of at the image of it. I can see everything at a vastly higher resolution. I can transcribe text that is so tiny in the image that it's invisible. I can even look around obstacles and see what's behind them if necessary. No LLM AI can do any of this.

    (1/6)

    #Long #LongPost #CWLong #CWLongPost #AltText #AltTextMeta #CWAltTextMeta #ImageDescription #ImageDescriptions #ImageDescriptionMeta #CWImageDescriptionMeta #AI #AIVsHuman #HumanVsAI
  15. @Georg Tuparev "The best image descriptions" as in better than other AI?

    Or as in describing all images better, at greater detail and with higher factual accuracy than any human possibly could, no exceptions? Even including human experts on an extreme niche topic?

    #ImageDescription #ImageDescriptions #ImageDescriptionMeta #CWImageDescriptionMeta #AI #AIVsHuman #HumanVsAI
  16. @nihilistic_capybara Yes. As a matter of fact, I've had an AI describe an image after describing it myself twice already. And I've always analysed the AI-generated description of the image from the point of view of someone who a) is very knowledgeable about these worlds in general and that very place in particular, b) has knowledge about the setting in the image which is not available anywhere on the Web because only he has this knowledge and c) can see much much more directly in-world than the AI can see in the scaled-down image.

    So here's an example.

    This was my first comparison thread. It may not look like it because it clearly isn't on Mastodon (at least I guess it's clear that this is not Mastodon), but it's still in the Fediverse, and it was sent to a whole number of Mastodon instances. Unfortunately, as I don't have any followers on layer8.space and didn't have any when I posted this, the post is not available on layer8.space. So you have to see it at the source in your Web browser rather than in your Mastodon app or otherwise on your Mastodon timeline.

    (Caution ahead: By my current standards, the image descriptions are outdated. Also, the explanations are not entirely accurate.)

    If you open the link, you'll see a post with a title, a summary and "View article" below. This works like Mastodon CWs because it's the exact same technology. Click or tap "View article" to see the full post. Warning: As the summary/CW indicates, it's very long.

    You'll see a bit of introduction post text, then the image with an alt-text that's actually short for my standards (on Mastodon, the image wouldn't be in the post, but below the post as a file attachment), then some more post text with the AI-generated image description and finally an additional long image description which is longer than 50 standard Mastodon toots. I've first used the same image, largely the same alt-text and the same long description in this post.

    Scroll further down, and you'll get to a comment in which I pick the AI description apart and analyse it for accuracy and detail level.

    For your convenience, here are some points where the AI failed:

    • The AI did not clearly identify the image as from a virtual world. It remained vague. Especially, it did not recognise the location as the central crossing at BlackWhite Castle in Pangea Grid, much less explain what either is. (Then again, explanations do not belong into alt-text. But when I posted the image, BlackWhite Castle had been online for two or three weeks and advertised on the Web for about as long.)
    • It failed to mention that the image is greyscale. That is, it actually failed to recognise that it isn't the image that's greyscale, but both the avatar and the entire scenery.
    • It referred to my avatar as a "character" and not an avatar.
    • It failed to recognise the avatar as my avatar.
    • It did not describe at all what my avatar looks like.
    • It hallucinated about what my avatar looks at. Allegedly, my avatar is looking at the advertising board towards the right. Actually, my avatar is looking at the cliff in the background which the AI does not mention at all. The AI could impossibly see my avatar's eyeballs from behind (and yes, they can move within the head).
    • It did not describe anything about the advertising board, especially not what's on it.
    • It did not know whether what it thinks my avatar is looking at is a sign or an information board, so it was still vague.
    • It hallucinated about a forest with a dense canopy. Actually, there are only a few trees, there is no canopy, the tops of the trees closer to the camera are not within the image, and the AI was confused by the mountain and the little bit of sky in the background.
    • The AI misjudged the lighting and hallucinated about the time of day, also because it doesn't know where the avatar and the camera are oriented.
    • It used the attributes "calm and serene" on something that's inspired by German black-and-white Edgar Wallace thrillers from the 1950s and the 1960s. It had no idea what's going on.
    • It did not mention a single bit of text in the image. Instead, it should have transcribed all of them verbatim. All of them. Legible in the image at the given resolution or not. (Granted, I myself forgot to transcribe a few little things in the image on the advertisement for the motel on the advertising board such as the license plate above the office door as well as the bits of text on the old map on the same board. But I didn't have any source for the map with a higher resolution, so I didn't give a detailed description of the map at all, and the text on it was illegible even to me.)
    • It did not mention that strange illuminated object towards the right at all. I'd expect a good AI to correctly identify it as an OpenSimWorld beacon, describe what it looks like, transcribe all text on it verbatim and, if asked for it, explain what it is, what it does and what it's there for in a way that everyone will understand. All 100% accurately.

    CC: @🅰🅻🅸🅲🅴  (🌈🦄)

    #Long #LongPost #CWLong #CWLongPost #OpenSim #OpenSimulator #Metaverse #VirtualWorlds #AltText #AltTextMeta #CWAltTextMeta #ImageDescription #ImageDescriptions #ImageDescriptionMeta #CWImageDescriptionMeta #AI #LLM #AIVsHuman #HumanVsAI
  17. @nihilistic_capybara LLMs aren't omniscient, and they will never be.

    If I make a picture on a sim in an OpenSim-based grid (that's a 3-D virtual world) which has only been started up for the first time 10 minutes ago, and which the WWW knows exactly zilch about, and I feed that picture to an LLM, I do not think the LLM will correctly pinpoint the place where the image was taken. It will not be able to correctly say that the picture was taken at <Place> on <Sim> in <Grid>, and then explain that <Grid> is a 3-D virtual world, a so-called grid, based on the virtual world server software OpenSimulator, and carry on explaining what OpenSim is, why a grid is called a grid, what a region is and what a sim is. But I can do that.

    If there's a sign with three lines of text on it somewhere within the borders of the image, but it's so tiny at the resolution of the image that it's only a few dozen pixels altogether, then no LLM will be able to correctly transcribe the three lines of text verbatim. It probably won't even be able to identify the sign as a sign. But I can do that by reading the sign not in the image, but directly in-world.

    By the way: All my original images are from within OpenSim grids. I've probably put more thought into describing images from virtual worlds than anyone. And I've pitted my own hand-written image description against an AI-generated image description of the self-same image twice. So I guess I know what I'm writing about.

    CC: @🅰🅻🅸🅲🅴  (🌈🦄) @nihilistic_capybara

    #Long #LongPost #CWLong #OpenSim #OpenSimulator #Metaverse #VirtualWorlds #CWLongPost #ImageDescription #ImageDescriptions #ImageDescriptionMeta #CWImageDescriptionMeta #AI #LLM #AIVsHuman #HumanVsAI
  18. @Justin Derrick The question, however, is: What is "high-quality"? How is it defined?

    Would the bot go by the definition valid for commercial/scientific/technological websites and blogs, i.e. ideally no more than 125 characters, and only a short and concise visual description with no further information?

    Or would the bot go by Mastodon's culture and Mastodon's standards, i.e. the longer and more detailed, the better, any and all extra information is welcome in alt-text (because it doesn't fit into the toot), and the limit is 1,500 characters?

    That is, if it were for me, the bot would go look both for alt-texts and for image descriptions in the post text body and judge both. Because I do both at the same time for my original images. An extremely detailed long image description in the post itself (character limit for post and alt-texts combined here: over 16 million) that also comes with all necessary explanations and transcripts of all text in the image, plus an alt-text that's as detailed as 1,500 characters (minus notification about the long description in the post) allow, but with no explanations, and I usually have to leave out text transcripts as well because they're too many.

    You may say the alt-text is superfluous if it's just a much shorter version of the long description. But as long as the Mastodon HOA demands there be an alt-text to every image, no matter what (especially seeing as I always hide my image posts behind summaries/content warnings, so you can't see right of the bat that there's a long image description in the post), I add alt-texts to my original images.

    I'm actually curious about how the bot would judge my descriptions. Maybe it'd flag them "inadequate" because it notices that the bits of text in the image are not transcribed in the alt-text. Maybe it'd be irritated because I have headlines in my long image descriptions, because they're so long that they need two levels of headlines. Maybe it'd flag them "inadequate" because it goes strictly by WCAG, and a) the alt-texts exceed 200 characters, b) long image descriptions do not belong into the text body by any known official accessibility standards, and c) neither my alt-texts nor my long descriptions are limited to what's supposed to be important within the context of the post.

    Anyway, in the meantime, you can follow the account @Alt Text Hall of Fame and the hashtag #AltTextHallOfFame.

    CC: @Simon Brooke

    #Long #LongPost #CWLong #CWLongPost #FediMeta #FediverseMeta #CWFediMeta #CWFediverseMeta #MastodonHOA #AltText #AltTextMeta #CWAltTextMeta #ImageDescription #ImageDescriptions #ImageDescriptionMeta #CWImageDescriptionMeta
  19. Jfc I got a lot of mileage out of one goddamn screenshot but hey!
    My b/Blind and vision-impaired or sight-loss fellows shall feast tonight!

    … If any of you care about niche multi-layered terminology for humorous agenerational interconnective co-created commentary on base genuine media, surprises, commentary, and intersections of such!

    You goddamn linguistic nerds!!

    Ooh, I feel like youtube.com/watch?v=YrHGbCyAqO is structurally related.
    OMG. THAT’S HOW ADHD BRAINS CONNECT IDEAS.

    youtube.com/watch?v=cQKGUgOfD8

    AuHD or AudHD (urgh, I hate that initialism) people connect ideas by their structural relevance (and personal ascription of importance).
    And, apparently, Neurotypical, neuro default, or neuro expected people connect ideas by narrative relevance! And, frankly, also how much importance they ascribe to the narrative, topic, or relevance.

    Omg I’m an edu blogger?! This is like finding out that non-binary gender was an option! Holy shit.

    #EaseOfAccess #Alt #ALT4you #ALT4me #alt4u #altText #imgDesc #ImageDescriptionMeta #ImageDescriptionPostMeta #omg #AGAIN #thisCallsFor #CrusherP #Crusher #Vocaloid #metaCommentary #IdidNotKnow #itWasVocaloid #forYears #blind #BlindFedi #BlindMasto #vision #VisionLoss #eyes #eyesight #visible #visibility #sensory #sensoryImpairment #sensoryLimit #sensoryLoad #sensoryOverload #information #InformationTechnology #InformationOverload #info #InfoTech #EduBlogger