Pop-Up Thingie

>>> Magnum BBS <<<
  • Home
  • Forum
  • Files
  • Log in

  1. Forum
  2. Usenet
  3. SCI.MISC
  • difficulty extracting data from PDFs

    From Retrograde@21:1/5 to All on Wed Mar 12 01:10:03 2025
    From the «cry me a river, AI» department:
    Title: Why Extracting Data from PDFs Remains a Nightmare for Data Experts Author: [email protected]
    Date: Tue, 11 Mar 2025 17:26:00 +0000
    Link: https://it.slashdot.org/story/25/03/11/1726218/why-extracting-data-from-pdfs-remains-a-nightmare-for-data-experts?utm_source=rss1.0mainlinkanon&utm_medium=feed

    Businesses, governments, and researchers continue to struggle with extracting usable data from PDF files, despite AI advances. These digital documents contain valuable information for everything from scientific research to government records, but their rigid formats make extraction difficult. "PDFs are a creature of a time when print layout was a big influence on publishing software," Derek Willis, a lecturer in Data and Computational Journalism at the University of Maryland, told ArsTechnica. This print-oriented design means many PDFs are essentially "pictures of information" requiring optical character recognition (OCR) technology. Traditional OCR systems have existed since the 1970s but struggle with complex layouts and poor-quality scans. New AI language models from companies like Google and Mistral now attempt to process documents more holistically, with varying success. "Right now, the clear leader is Google's Gemini 2.0 Flash Pro Experimental," Willis notes, while Mistral's recent OCR solution "performed poorly" in tests.

    [image 2][2][image 4][4]

    Read more of this story[5] at Slashdot.

    Links:
    [1]: http://twitter.com/home?status=Why+Extracting+Data+from+PDFs+Remains+a+Nightmare+for+Data+Experts%3A+https%3A%2F%2Fit.slashdot.org%2Fstory%2F25%2F03%2F11%2F1726218%2F%3Futm_source%3Dtwitter%26utm_medium%3Dtwitter (link)
    [2]: https://a.fsdn.com/sd/twitter_icon_large.png (image)
    [3]: http://www.facebook.com/sharer.php?u=https%3A%2F%2Fit.slashdot.org%2Fstory%2F25%2F03%2F11%2F1726218%2Fwhy-extracting-data-from-pdfs-remains-a-nightmare-for-data-experts%3Futm_source%3Dslashdot%26utm_medium%3Dfacebook (link)
    [4]: https://a.fsdn.com/sd/facebook_icon_large.png (image)
    [5]: https://it.slashdot.org/story/25/03/11/1726218/why-extracting-data-from-pdfs-remains-a-nightmare-for-data-experts?utm_source=rss1.0moreanon&utm_medium=feed (link)

    --- SoupGate-Win32 v1.05
    * Origin: fsxNet Usenet Gateway (21:1/5)
  • From anthk@21:1/5 to Retrograde on Tue Mar 18 11:23:39 2025
    On 2025-03-12, Retrograde <[email protected]d> wrote:
    From the «cry me a river, AI» department:
    Title: Why Extracting Data from PDFs Remains a Nightmare for Data Experts Author: [email protected]
    Date: Tue, 11 Mar 2025 17:26:00 +0000
    Link: https://it.slashdot.org/story/25/03/11/1726218/why-extracting-data-from-pdfs-remains-a-nightmare-for-data-experts?utm_source=rss1.0mainlinkanon&utm_medium=feed

    Businesses, governments, and researchers continue to struggle with extracting usable data from PDF files, despite AI advances. These digital documents contain valuable information for everything from scientific research to government records, but their rigid formats make extraction difficult. "PDFs are a creature of a time when print layout was a big influence on publishing software," Derek Willis, a lecturer in Data and Computational Journalism at the
    University of Maryland, told ArsTechnica. This print-oriented design means many
    PDFs are essentially "pictures of information" requiring optical character recognition (OCR) technology. Traditional OCR systems have existed since the 1970s but struggle with complex layouts and poor-quality scans. New AI language
    models from companies like Google and Mistral now attempt to process documents
    more holistically, with varying success. "Right now, the clear leader is Google's Gemini 2.0 Flash Pro Experimental," Willis notes, while Mistral's recent OCR solution "performed poorly" in tests.

    [image 2][2][image 4][4]

    Read more of this story[5] at Slashdot.

    Links:
    [1]: http://twitter.com/home?status=Why+Extracting+Data+from+PDFs+Remains+a+Nightmare+for+Data+Experts%3A+https%3A%2F%2Fit.slashdot.org%2Fstory%2F25%2F03%2F11%2F1726218%2F%3Futm_source%3Dtwitter%26utm_medium%3Dtwitter (link)
    [2]: https://a.fsdn.com/sd/twitter_icon_large.png (image)
    [3]: http://www.facebook.com/sharer.php?u=https%3A%2F%2Fit.slashdot.org%2Fstory%2F25%2F03%2F11%2F1726218%2Fwhy-extracting-data-from-pdfs-remains-a-nightmare-for-data-experts%3Futm_source%3Dslashdot%26utm_medium%3Dfacebook (link)
    [4]: https://a.fsdn.com/sd/facebook_icon_large.png (image)
    [5]: https://it.slashdot.org/story/25/03/11/1726218/why-extracting-data-from-pdfs-remains-a-nightmare-for-data-experts?utm_source=rss1.0moreanon&utm_medium=feed (link)

    Why not Recoll under Linux/Unix/Mac/Windows?

    https://www.recoll.org/index.html

    Recoll, not Recall.

    --- SoupGate-Win32 v1.05
    * Origin: fsxNet Usenet Gateway (21:1/5)
  • Who's Online

  • Recent Visitors

    • Rixter
      Wed Jul 29 02:00:40 2026
      from Madison, Nc via Telnet
    • Centurion
      Tue Jul 28 22:54:59 2026
      from Berea, Ohio via Telnet
    • Bob Worm
      Tue Jul 28 16:01:18 2026
      from Wales, Uk via Telnet
    • Rixter
      Tue Jul 28 13:42:46 2026
      from Madison, Nc via Telnet
    • Krenn
      Tue Jul 28 11:59:57 2026
      from Sydney, Nsw via Telnet
    • Rixter
      Tue Jul 28 01:23:48 2026
      from Madison, Nc via Telnet
    • Centurion
      Mon Jul 27 22:50:42 2026
      from Berea, Ohio via Telnet
    • Ataricrypt
      Mon Jul 27 19:19:17 2026
      from England via Telnet
  • System Info

    Sysop: Keyop
    Location: Huddersfield, West Yorkshire, UK
    Users: 741
    Nodes: 16 (2 / 14)
    Uptime: 60:35:46
    Calls: 12,446
    Calls today: 1
    Files: 15,192
    Messages: 6,537,450

© >>> Magnum BBS <<<, 2026