Skip to main content
BRILLIQS

Data Engineering

Processing engines, streaming platforms and lakehouses that move data reliably at scale.

27 tools

Apache SparkData Engineering

Apache Spark

An engine for processing data across many machines, using the same code whether the data is small or large.

Explore
Apache KafkaData Engineering

Apache Kafka

A distributed log where records are appended and kept, so many consumers can read them at their own pace.

Explore
Apache AirflowData Engineering

Apache Airflow

An open source platform for building, scheduling and monitoring batch data workflows written in Python.

Explore
Azure Synapse AnalyticsData Modernization

Azure Synapse Analytics

An Azure service bringing warehouse queries, Spark processing and data integration together in one workspace.

Explore
Apache FlinkData Engineering

Apache Flink

A framework for processing continuous data streams with local state, event time and exactly once recovery.

Explore
Apache NiFiData Engineering

Apache NiFi

A platform for automating the flow of data between systems, built and controlled from a visual interface.

Explore
Apache BeamData Engineering

Apache Beam

One programming model for batch and streaming pipelines that can run on several different processing engines.

Explore
Apache HadoopData Engineering

Apache Hadoop

A framework for storing and processing large datasets across a cluster of machines.

Explore
Apache HiveData Engineering

Apache Hive

A data warehouse system that lets you query very large files in distributed storage using SQL.

Explore
Apache IcebergData Engineering

Apache Iceberg

An open table format that lets several engines read and write the same large analytic tables safely.

Explore
Apache HudiData Engineering

Apache Hudi

A data lake platform built around record keys, so individual rows can be updated and changes can be read incrementally.

Explore
Apache PulsarData Engineering

Apache Pulsar

A messaging and streaming platform that keeps its serving layer and its storage layer separate.

Explore
Amazon KinesisData Engineering

Amazon Kinesis

An AWS service for collecting streaming records continuously and making them available to consumer applications.

Explore
AirbyteData Engineering

Airbyte

An open source platform for moving data between sources and destinations, with connectors you can also build yourself.

Explore
AlteryxData Analytics

Alteryx

A tool where data preparation and analysis are built by dragging tools onto a canvas and connecting them into a workflow.

Explore
AWS GlueData Engineering

AWS Glue

A serverless AWS service that catalogues your data and runs the jobs that prepare it, without a cluster to manage.

Explore
Azure Data FactoryData Engineering

Azure Data Factory

A cloud service for building pipelines that move data between systems and orchestrate the steps around that movement.

Explore
ClouderaData Modernization

Cloudera

A commercial data platform assembling Hadoop ecosystem projects with management, security and support around them.

Explore
Amazon EMRData Engineering

Amazon EMR

An AWS platform for running open source big data frameworks such as Spark, Hive and Presto on managed clusters.

Explore
Apache DruidData Engineering

Apache Druid

A database built for fast analytical queries over event data, where time is treated as a first class column.

Explore
ClickHouseData Engineering

ClickHouse

A column oriented database management system that answers analytical SQL queries over very large tables.

Explore
Apache StormData Engineering

Apache Storm

A distributed system for processing unbounded streams of records as they arrive, one record at a time.

Explore
Apache ArrowData Engineering

Apache Arrow

A standard way of laying out columnar data in memory so different tools and languages can share it without converting.

Explore
Apache SqoopData Engineering

Apache Sqoop

A command line tool for bulk transfer of data between relational databases and Hadoop storage.

Explore
Apache FlumeData Engineering

Apache Flume

A service for collecting log data from many machines and delivering it to central storage through configured agents.

Explore
Apache OozieData Engineering

Apache Oozie

A workflow scheduler for Hadoop jobs, where workflows are defined in XML and can wait for data before running.

Explore
Apache ZooKeeperData Engineering

Apache ZooKeeper

A coordination service that distributed systems use to agree on shared state such as configuration and leadership.

Explore
All tools

Working with data engineering?

Our engineers can assess your current setup and tell you what is worth changing.

Book a Consultation