Apache Spark
An engine for processing data across many machines, using the same code whether the data is small or large.
ExploreApache Kafka
A distributed log where records are appended and kept, so many consumers can read them at their own pace.
ExploreApache Airflow
An open source platform for building, scheduling and monitoring batch data workflows written in Python.
ExploreAzure Synapse Analytics
An Azure service bringing warehouse queries, Spark processing and data integration together in one workspace.
ExploreApache Flink
A framework for processing continuous data streams with local state, event time and exactly once recovery.
ExploreApache NiFi
A platform for automating the flow of data between systems, built and controlled from a visual interface.
ExploreApache Beam
One programming model for batch and streaming pipelines that can run on several different processing engines.
ExploreApache Hadoop
A framework for storing and processing large datasets across a cluster of machines.
ExploreApache Hive
A data warehouse system that lets you query very large files in distributed storage using SQL.
ExploreApache Iceberg
An open table format that lets several engines read and write the same large analytic tables safely.
ExploreApache Hudi
A data lake platform built around record keys, so individual rows can be updated and changes can be read incrementally.
ExploreApache Pulsar
A messaging and streaming platform that keeps its serving layer and its storage layer separate.
ExploreAmazon Kinesis
An AWS service for collecting streaming records continuously and making them available to consumer applications.
ExploreAirbyte
An open source platform for moving data between sources and destinations, with connectors you can also build yourself.
ExploreAlteryx
A tool where data preparation and analysis are built by dragging tools onto a canvas and connecting them into a workflow.
ExploreAWS Glue
A serverless AWS service that catalogues your data and runs the jobs that prepare it, without a cluster to manage.
ExploreAzure Data Factory
A cloud service for building pipelines that move data between systems and orchestrate the steps around that movement.
ExploreCloudera
A commercial data platform assembling Hadoop ecosystem projects with management, security and support around them.
ExploreAmazon EMR
An AWS platform for running open source big data frameworks such as Spark, Hive and Presto on managed clusters.
ExploreApache Druid
A database built for fast analytical queries over event data, where time is treated as a first class column.
ExploreClickHouse
A column oriented database management system that answers analytical SQL queries over very large tables.
ExploreApache Storm
A distributed system for processing unbounded streams of records as they arrive, one record at a time.
ExploreApache Arrow
A standard way of laying out columnar data in memory so different tools and languages can share it without converting.
ExploreApache Sqoop
A command line tool for bulk transfer of data between relational databases and Hadoop storage.
ExploreApache Flume
A service for collecting log data from many machines and delivering it to central storage through configured agents.
ExploreApache Oozie
A workflow scheduler for Hadoop jobs, where workflows are defined in XML and can wait for data before running.
ExploreApache ZooKeeper
A coordination service that distributed systems use to agree on shared state such as configuration and leadership.
ExploreWorking with data engineering?
Our engineers can assess your current setup and tell you what is worth changing.