Awesome ClickHouse › ETL and Data Processing
wgzhao/Addax
⭐ 1436
Java
repository created 2019-07-17
Addax is a fast, versatile, and extensible open-source ETL (Extract, Transform, Load) tool designed to facilitate seamless data transfer between a wide variety of SQL and NoSQL data sources. It originated as a fork and evolution of Alibaba's DataX, enhancing its architecture and functionality to support over 20 different data sources including relational databases, NoSQL databases, and file systems. Notably, Addax supports ClickHouse among many other platforms such as Cassandra, MongoDB, MySQL, Oracle, PostgreSQL, Elasticsearch, Redis, and more, making it a comprehensive solution for data integration needs.
The tool is configured through simple JSON-based job descriptions, allowing users to define data transfer tasks with ease and flexibility. It offers a growing ecosystem of plugins that can be extended to support additional data sources or customized transformations, catering to diverse ETL scenarios. Addax is actively maintained, with improvements in performance, usability, and deployment options including Docker images for quick and convenient setup.
Addax provides detailed documentation and examples to help users get started quickly, including installation via Docker, installation scripts, or compiling from source. It supports Java Runtime (JDK 1.8+) and Python (for Windows environments), ensuring compatibility with common development environments. The project also emphasizes code quality and maintainability, following Java coding conventions and providing guidelines for contributors.
Overall, Addax is a powerful ETL tool suitable for enterprises and developers looking to integrate data across heterogeneous systems efficiently. Its support for ClickHouse and many other data sources, combined with its extensibility and ease of use, makes it a valuable asset for data engineering workflows and big data processing pipelines.
https://github.com/wgzhao/Addax
big-dataclickhousedata-engineeringdata-integrationdata-sourcesdata-transferdatabasedataxdockeretlexcelextensiblehadoophdfshiveimpalainfluxdbjavajson-configurationkudumysqlnosqlopen-sourceoraclepluginspostgresqlsqlsqlservertrino
Also in ETL and Data Processing
PeerDB is a high-performance, PostgreSQL-optimized ETL tool that enables fast, reliable, and cost-effective streaming of data from Postgres to data warehouses, queues, and storage engines, with native integration in ClickHouse Cloud.
YTsaurus is a scalable, fault-tolerant open-source big data platform featuring MapReduce, SQL engine, NoSQL store, and integration with ClickHouse for fast analytics.
Trench is an open-source, production-ready analytics infrastructure built on ClickHouse and Kafka for scalable, real-time event tracking and analytics with GDPR compliance.
Gluten is a middle layer that offloads JVM-based SQL engines' execution, such as Spark SQL, to high-performance native engines like ClickHouse and Velox, leveraging vectorized processing for accele...
DungBeetle is a distributed job server for asynchronously queuing and executing heavy SQL read jobs on MySQL, PostgreSQL, and ClickHouse databases, designed to offload report generation and improve application performance.
ClickBench is a comprehensive and reproducible benchmark designed to evaluate the performance of analytical databases, including ClickHouse, using realistic workloads derived from real-world web analytics data.
DataCap is an integrated software platform for data transformation, integration, and visualization, supporting a wide range of data sources including ClickHouse and other major databases.
VulcanSQL is an Analytical Data API Framework that simplifies and accelerates the creation of secure, scalable RESTful APIs from databases and data warehouses for AI agents and data applications.