Menu
Hackathon geared toward the 'liberation' of data from public PDF documents

Hackathon geared toward the 'liberation' of data from public PDF documents

The Sunlight Foundation and others will sponsor a three-day hackathon starting Friday

Massive amounts of unstructured data are held in the form of PDF documents, but extracting key figures and words out of PDFs in a programmatic manner can be difficult and costly. This poses a challenge to public-interest groups, journalists and others who are interested in running large-scale analyses on PDF documents in order to uncover valuable insights.

In a hackathon set for this week, participants will work on ways to improve the open-source software tools available for PDF data extraction.

"Say, for example, you want to model student loan securitizations," wrote Marc Joffe, principal consultant at Public Sector Credit Solutions and an organizer along with the Sunlight Foundation and others of the PDF Liberation Hackathon, in a guest post on the Mathbabe blog. "A corporation or well funded research institution can purchase an expensive, enterprise-level ETL (Extract-Transform-Load) tool to migrate data from the PDFs into a database. But this is not much help to insurgent modelers who want to produce open source work."

"Data journalists face a similar challenge," he added. "They often need to extract bulk data from PDFs to support their reporting. Examples include IRS Form 990s filed by non-profits and budgets issued by governments at all levels."

Data journalists have developed open-source PDF harvesting tools such as Tabula, Joffe added.

"Unfortunately, the free and low cost tools available to modelers, data journalists and transparency advocates have limitations that hinder their ability to handle large scale tasks," he wrote. "If, like me, you want to submit hundreds of PDFs to a software tool, press 'Go' and see large volumes of cleanly formatted data, you are out of luck."

The hackathon runs from Friday through Sunday and will be held at six sites, including the Sunlight Foundation's headquarters in Washington, D.C., according to the event's website. Remote participation is also possible.

Contestants will be able to work on "a PDF extraction challenge provided by one of our sponsoring organizations, can work on their own challenges or develop enhancements to an open source PDF extraction tool," according to the site.

While the use of open-source tools is encouraged, commercial tools are allowed as long as licensing costs less than US$1,000 and an unlimited trial is available.

It's true that some of the best tools for PDF extraction are proprietary and expensive, said analyst Curt Monash of Monash Research, who closely tracks the database and data-analysis market as well as public policy on technology.

"One of the leading filter/extraction libraries was bought by Verity, which was bought by Autonomy, which was bought by HP," he said via email on Thursday. "Another one, with a somewhat different orientation, was developed by Xerox, which spun it out as Inxight, which was bought by Business Objects, which was bought by SAP."

"It's worth remembering that there's a multi-stage process here," Monash added. "For example, a PDF can be converted to text (and image) data, (Name, value) pairs can be extracted. Those can have their spelling corrected. Then the company names can be regularized. In real life, there can be tens of steps."

As for the hackathon's potential value, "a large fraction of the world's interesting information is on paper, or in paper-like formats such as PDF," he added. "Of course it's worthwhile to make all that more accessible."

Chris Kanaracus covers enterprise software and general technology breaking news for The IDG News Service. Chris' email address is Chris_Kanaracus@idg.com


Follow Us

Join the newsletter!

Or

Sign up to gain exclusive access to email subscriptions, event invitations, competitions, giveaways, and much more.

Membership is free, and your security and privacy remain protected. View our privacy policy before signing up.

Error: Please check your email address.

Tags privacyinternetlegalsoftwareapplicationsapplication developmentSunlight Foundation

Featured

Slideshows

Reseller News kicks off awards season in 2019 with Judges' Lunch

Reseller News kicks off awards season in 2019 with Judges' Lunch

The 2019 Reseller News Innovation Awards has kicked off with the Judges Lunch in Auckland with 70 judges in the voting panel. The awards will reflect the changing dynamics of the channel, recognising excellence across customer value and innovation - spanning start-ups, partners, distributors and vendors. Photos by Christine Wong.

Reseller News kicks off awards season in 2019 with Judges' Lunch
Reseller News welcomes industry figures for 2019 Hall of Fame lunch

Reseller News welcomes industry figures for 2019 Hall of Fame lunch

Reseller News welcomed 2018 inductees - Chris Simpson, Kendra Ross and Phill Patton - to the third running of the Reseller News Hall of Fame lunch, held at the French Cafe in Auckland. The inductees discussed the changing landscape of the technology industry in New Zealand, while outlining ways to attract a new breed of players to the ecosystem. Photos by Gino Demeer.

Reseller News welcomes industry figures for 2019 Hall of Fame lunch
Upcoming tech talent share insights at inaugural Emerging Leaders Forum 2019

Upcoming tech talent share insights at inaugural Emerging Leaders Forum 2019

The channel came together for the inaugural Reseller News Emerging Leaders Forum in New Zealand, created to provide a program that identifies, educates and showcases the upcoming talent of the ICT industry. Hosted as a half day event, attendees heard from industry champions as keynoters and panelists talked about future opportunities and leadership paths and joined mentoring sessions with members of the ICT industry Hall of Fame. The forum concluded with 30 Under 30 Tech Awards across areas of Sales, Entrepreneur, Marketing, Management, Technical and Human Resources. Photos by Gino Demeer.

Upcoming tech talent share insights at inaugural Emerging Leaders Forum 2019
Show Comments