#pymupdf4llm — Public Fediverse posts
Live and recent posts from across the Fediverse tagged #pymupdf4llm, aggregated by home.social.
-
Невидимый слой PDF с инструкциями для ботов
Невидимый текстовый слой PDF можно редактировать и экспортировать в Markdown, JSON и TXT. Такие документы называются адаптивными PDF , они созданы для чтения и людьми, и роботами. Люди видят обычный PDF, а роботы — отдельный слой ActualText с текстом в Markdown и картинками в base64.
https://habr.com/ru/companies/globalsign/articles/1058328/
#markdown #разметка #PyMuPDF #умный_PDF #Adaptivepdf #PyMuPDF4LLM #OCR #pdx #Poppler
-
Невидимый слой PDF с инструкциями для ботов
Невидимый текстовый слой PDF можно редактировать и экспортировать в Markdown, JSON и TXT. Такие документы называются адаптивными PDF , они созданы для чтения и людьми, и роботами. Люди видят обычный PDF, а роботы — отдельный слой ActualText с текстом в Markdown и картинками в base64.
https://habr.com/ru/companies/globalsign/articles/1058328/
#markdown #разметка #PyMuPDF #умный_PDF #Adaptivepdf #PyMuPDF4LLM #OCR #pdx #Poppler
-
Невидимый слой PDF с инструкциями для ботов
Невидимый текстовый слой PDF можно редактировать и экспортировать в Markdown, JSON и TXT. Такие документы называются адаптивными PDF , они созданы для чтения и людьми, и роботами. Люди видят обычный PDF, а роботы — отдельный слой ActualText с текстом в Markdown и картинками в base64.
https://habr.com/ru/companies/globalsign/articles/1058328/
#markdown #разметка #PyMuPDF #умный_PDF #Adaptivepdf #PyMuPDF4LLM #OCR #pdx #Poppler
-
🔍 #PymuPDF4llm revolutionizes #PDFOfTheDay data extraction for #LLM applications:
📋 Extracts structured text, tables, and images from PDFs, converting them to clean markdown format for optimal #AI processing
🛠️ Features include:
- Word-by-word extraction capability
- Custom image format & DPI settings
- Table conversion to CSV/JSON
- Page chunking options💻 Key technical benefits:
- Simple pip installation
- #Python integration
- Full #opensource availability
- No usage limitations or credits system🔧 Perfect for:
- #DataScience projects
- Document processing pipelines
- #AI training data preparation
- Automated workflow systemsSource: https://github.com/deepset-ai/pymupdf4llm
https://pypi.org/project/pymupdf4llm/ -
🔍 #PymuPDF4llm revolutionizes #PDFOfTheDay data extraction for #LLM applications:
📋 Extracts structured text, tables, and images from PDFs, converting them to clean markdown format for optimal #AI processing
🛠️ Features include:
- Word-by-word extraction capability
- Custom image format & DPI settings
- Table conversion to CSV/JSON
- Page chunking options💻 Key technical benefits:
- Simple pip installation
- #Python integration
- Full #opensource availability
- No usage limitations or credits system🔧 Perfect for:
- #DataScience projects
- Document processing pipelines
- #AI training data preparation
- Automated workflow systemsSource: https://github.com/deepset-ai/pymupdf4llm
https://pypi.org/project/pymupdf4llm/ -
🔍 #PymuPDF4llm revolutionizes #PDFOfTheDay data extraction for #LLM applications:
📋 Extracts structured text, tables, and images from PDFs, converting them to clean markdown format for optimal #AI processing
🛠️ Features include:
- Word-by-word extraction capability
- Custom image format & DPI settings
- Table conversion to CSV/JSON
- Page chunking options💻 Key technical benefits:
- Simple pip installation
- #Python integration
- Full #opensource availability
- No usage limitations or credits system🔧 Perfect for:
- #DataScience projects
- Document processing pipelines
- #AI training data preparation
- Automated workflow systemsSource: https://github.com/deepset-ai/pymupdf4llm
https://pypi.org/project/pymupdf4llm/ -
🔍 #PymuPDF4llm revolutionizes #PDFOfTheDay data extraction for #LLM applications:
📋 Extracts structured text, tables, and images from PDFs, converting them to clean markdown format for optimal #AI processing
🛠️ Features include:
- Word-by-word extraction capability
- Custom image format & DPI settings
- Table conversion to CSV/JSON
- Page chunking options💻 Key technical benefits:
- Simple pip installation
- #Python integration
- Full #opensource availability
- No usage limitations or credits system🔧 Perfect for:
- #DataScience projects
- Document processing pipelines
- #AI training data preparation
- Automated workflow systemsSource: https://github.com/deepset-ai/pymupdf4llm
https://pypi.org/project/pymupdf4llm/ -
🔍 #PymuPDF4llm revolutionizes #PDFOfTheDay data extraction for #LLM applications:
📋 Extracts structured text, tables, and images from PDFs, converting them to clean markdown format for optimal #AI processing
🛠️ Features include:
- Word-by-word extraction capability
- Custom image format & DPI settings
- Table conversion to CSV/JSON
- Page chunking options💻 Key technical benefits:
- Simple pip installation
- #Python integration
- Full #opensource availability
- No usage limitations or credits system🔧 Perfect for:
- #DataScience projects
- Document processing pipelines
- #AI training data preparation
- Automated workflow systemsSource: https://github.com/deepset-ai/pymupdf4llm
https://pypi.org/project/pymupdf4llm/ -
RAG/LLMの前処理:PyMuPDF4LLMを使用してPDFをMarkdownへ変換する
https://qiita.com/cyberBOSE/items/c276d273bfc20881adfc?utm_campaign=popular_items&utm_medium=feed&utm_source=popular_items -
RAG/LLMの前処理:PyMuPDF4LLMを使用してPDFをMarkdownへ変換する
https://qiita.com/cyberBOSE/items/c276d273bfc20881adfc?utm_campaign=popular_items&utm_medium=feed&utm_source=popular_items -
RAG/LLMの前処理:PyMuPDF4LLMを使用してPDFをMarkdownへ変換する
https://qiita.com/cyberBOSE/items/c276d273bfc20881adfc?utm_campaign=popular_items&utm_medium=feed&utm_source=popular_items -
RAG/LLMの前処理:PyMuPDF4LLMを使用してPDFをMarkdownへ変換する
https://qiita.com/cyberBOSE/items/c276d273bfc20881adfc?utm_campaign=popular_items&utm_medium=feed&utm_source=popular_items