home.social

#pymupdf4llm — Public Fediverse posts

Live and recent posts from across the Fediverse tagged #pymupdf4llm, aggregated by home.social.

fetched live
  1. Невидимый слой PDF с инструкциями для ботов

    Невидимый текстовый слой PDF можно редактировать и экспортировать в Markdown, JSON и TXT. Такие документы называются адаптивными PDF , они созданы для чтения и людьми, и роботами. Люди видят обычный PDF, а роботы — отдельный слой ActualText с текстом в Markdown и картинками в base64.

    habr.com/ru/companies/globalsi

    #markdown #разметка #PyMuPDF #умный_PDF #Adaptivepdf #PyMuPDF4LLM #OCR #pdx #Poppler

  2. Невидимый слой PDF с инструкциями для ботов

    Невидимый текстовый слой PDF можно редактировать и экспортировать в Markdown, JSON и TXT. Такие документы называются адаптивными PDF , они созданы для чтения и людьми, и роботами. Люди видят обычный PDF, а роботы — отдельный слой ActualText с текстом в Markdown и картинками в base64.

    habr.com/ru/companies/globalsi

    #markdown #разметка #PyMuPDF #умный_PDF #Adaptivepdf #PyMuPDF4LLM #OCR #pdx #Poppler

  3. Невидимый слой PDF с инструкциями для ботов

    Невидимый текстовый слой PDF можно редактировать и экспортировать в Markdown, JSON и TXT. Такие документы называются адаптивными PDF , они созданы для чтения и людьми, и роботами. Люди видят обычный PDF, а роботы — отдельный слой ActualText с текстом в Markdown и картинками в base64.

    habr.com/ru/companies/globalsi

    #markdown #разметка #PyMuPDF #умный_PDF #Adaptivepdf #PyMuPDF4LLM #OCR #pdx #Poppler

  4. 🔍 #PymuPDF4llm revolutionizes #PDFOfTheDay data extraction for #LLM applications:

    📋 Extracts structured text, tables, and images from PDFs, converting them to clean markdown format for optimal #AI processing

    🛠️ Features include:
    - Word-by-word extraction capability
    - Custom image format & DPI settings
    - Table conversion to CSV/JSON
    - Page chunking options

    💻 Key technical benefits:
    - Simple pip installation
    - #Python integration
    - Full #opensource availability
    - No usage limitations or credits system

    🔧 Perfect for:
    - #DataScience projects
    - Document processing pipelines
    - #AI training data preparation
    - Automated workflow systems

    Source: github.com/deepset-ai/pymupdf4
    pypi.org/project/pymupdf4llm/

  5. 🔍 #PymuPDF4llm revolutionizes #PDFOfTheDay data extraction for #LLM applications:

    📋 Extracts structured text, tables, and images from PDFs, converting them to clean markdown format for optimal #AI processing

    🛠️ Features include:
    - Word-by-word extraction capability
    - Custom image format & DPI settings
    - Table conversion to CSV/JSON
    - Page chunking options

    💻 Key technical benefits:
    - Simple pip installation
    - #Python integration
    - Full #opensource availability
    - No usage limitations or credits system

    🔧 Perfect for:
    - #DataScience projects
    - Document processing pipelines
    - #AI training data preparation
    - Automated workflow systems

    Source: github.com/deepset-ai/pymupdf4
    pypi.org/project/pymupdf4llm/

  6. 🔍 #PymuPDF4llm revolutionizes #PDFOfTheDay data extraction for #LLM applications:

    📋 Extracts structured text, tables, and images from PDFs, converting them to clean markdown format for optimal #AI processing

    🛠️ Features include:
    - Word-by-word extraction capability
    - Custom image format & DPI settings
    - Table conversion to CSV/JSON
    - Page chunking options

    💻 Key technical benefits:
    - Simple pip installation
    - #Python integration
    - Full #opensource availability
    - No usage limitations or credits system

    🔧 Perfect for:
    - #DataScience projects
    - Document processing pipelines
    - #AI training data preparation
    - Automated workflow systems

    Source: github.com/deepset-ai/pymupdf4
    pypi.org/project/pymupdf4llm/

  7. 🔍 #PymuPDF4llm revolutionizes #PDFOfTheDay data extraction for #LLM applications:

    📋 Extracts structured text, tables, and images from PDFs, converting them to clean markdown format for optimal #AI processing

    🛠️ Features include:
    - Word-by-word extraction capability
    - Custom image format & DPI settings
    - Table conversion to CSV/JSON
    - Page chunking options

    💻 Key technical benefits:
    - Simple pip installation
    - #Python integration
    - Full #opensource availability
    - No usage limitations or credits system

    🔧 Perfect for:
    - #DataScience projects
    - Document processing pipelines
    - #AI training data preparation
    - Automated workflow systems

    Source: github.com/deepset-ai/pymupdf4
    pypi.org/project/pymupdf4llm/

  8. 🔍 #PymuPDF4llm revolutionizes #PDFOfTheDay data extraction for #LLM applications:

    📋 Extracts structured text, tables, and images from PDFs, converting them to clean markdown format for optimal #AI processing

    🛠️ Features include:
    - Word-by-word extraction capability
    - Custom image format & DPI settings
    - Table conversion to CSV/JSON
    - Page chunking options

    💻 Key technical benefits:
    - Simple pip installation
    - #Python integration
    - Full #opensource availability
    - No usage limitations or credits system

    🔧 Perfect for:
    - #DataScience projects
    - Document processing pipelines
    - #AI training data preparation
    - Automated workflow systems

    Source: github.com/deepset-ai/pymupdf4
    pypi.org/project/pymupdf4llm/