Menu
Facebook query software crawls data mountains more quickly

Facebook query software crawls data mountains more quickly

Facebook has updated its Presto SQL query engine for faster performance

Facebook engineers have revamped the company's open source Presto query engine to run up to four times faster, making it even more feasible for very large scale data warehouse work.

"We have a large number of internal users at Facebook who use Presto on a continuous basis for data analysis. Improving query performance directly improves their productivity, so we thought through ways to make Presto even faster," wrote Dain Sundstrom, in a blog post introducing the improvements.

Facebook created Presto to run SQL commands against large amounts of unstructured data, notably data the company kept in Hadoop file systems. Understood by countless database administrators, SQL (Structured Query Language) has long been the cornerstone of relational database systems and data analysis systems.

Last November, Facebook open sourced Presto. Like with many of the programs the company has developed in-house, Facebook is hoping others will deploy the code and submit bug fixes and improvements. Organizations could use the software to run data analysis on sets of data that would be too large, or too expensive to fit into a commercial data warehouse.

To improve performance, Facebook engineers revamped Presto's software for reading data. The software can now directly ingest data from Hadoop systems in columns, rather than by reading them in by rows, which required additional time for restructuring the data.

The software now takes note of the minimum and maximum values of each column, which allows it to more quickly find a set of data with a small range of values. It also evaluates queries more intelligently, employing the user's filter terms to minimize the data set being inspected, saving processing time.

With these improvements, Facebook engineers found that Presto was able to execute single single column reads up to four times faster, as measured against a 600-million-row data set residing on 14 servers, with each server running 16 cores and 64GB of memory.

Presto seems to be gaining popularity outside of Facebook. Argyle uses the software as the basis of its own real-time fraud detection software for telecommunications companies. Presto "gives us a lot of flexibility," said Arshak Navruzyan, Argyle vice president of product management.

Argyle paired Presto with open source data store Accumolo. Argyle's customers keep terabytes, and even petabytes worth of user call data, which is far too large to hold in traditional data warehouses.

Presto queries serve as the basis of Argyle's alert messaging system, which flags unusual behavior for its customers. For Argyle, Presto provides an interface for user-facing business intelligence software such as Tableau.

Airbnb also uses Presto, in conjunction with its own software, to allow employees to access the company's operational data.

Presto is one of a number of SQL query engines developed for interrogating massive amounts of data across multiple servers in parallel. Cloudera developed Impala for similar reasons. Pivotal also created HAWQ for the same purpose, borrowing technology from its Greenplum database technology.


Follow Us

Join the New Zealand Reseller News newsletter!

Error: Please check your email address.

Tags softwareFacebook

Featured

Slideshows

Kiwi channel comes together for another round of After Hours

Kiwi channel comes together for another round of After Hours

The channel came together for another round of After Hours, with a bumper crowd of distributors, vendors and partners descending on The Jefferson in Auckland. Photos by Maria Stefina.​

Kiwi channel comes together for another round of After Hours
Consegna comes to town with AWS cloud offerings launch in Auckland

Consegna comes to town with AWS cloud offerings launch in Auckland

Emerging start-up Consegna has officially launched its cloud offerings in the New Zealand market, through a kick-off event held at Seafarers Building in Auckland.​ Founded in June 2016, the Auckland-based business is backed by AWS and supported by a global team of cloud specialists, leveraging global managed services partnerships with Rackspace locally.

Consegna comes to town with AWS cloud offerings launch in Auckland
Veritas honours top performing trans-Tasman partners

Veritas honours top performing trans-Tasman partners

Veritas honoured its top performing partners across the channel in Australia and New Zealand, recognising innovation and excellence on both sides of the Tasman. Revealed under the Vivid lights in Sydney, Intalock claimed the coveted Partner of the Year 2017 (Pacific) award, with Data#3 acknowledged for 12 months of strong growth across the market. Meanwhile, Datacom took home the New Zealand honours, with Global Storage and Insentra winning service provider and consulting awards respectively. Dicker Data was recognised as the standout distributor of the year, while Hitachi Data Systems claimed the alliance partner award. Photos by Bob Seary.

Veritas honours top performing trans-Tasman partners
Show Comments