Awesome ClickHouseETL and Data Processing

devlive-community/datacap

⭐ 1056 Java repository created 2022-09-17

DataCap is an integrated software platform designed for comprehensive data transformation, integration, and visualization. It supports a wide variety of data sources including file types, big data-related databases, relational databases, and NoSQL databases. The software enables users to manage multiple data sources efficiently and perform various data operations and conversions. It also provides capabilities for creating data charts and monitoring data sources, making it a versatile tool for data management and analysis. The platform supports querying data from any SQL-speaking datastore or data engine, with explicit support for ClickHouse, MySQL, Presto, and many other major database solutions. This includes a broad range of databases such as Redis, PostgreSQL, Hive, Elasticsearch, MongoDB, Oracle, SQL Server, and many more, covering both traditional and modern data storage technologies. This extensive connector support allows DataCap to integrate seamlessly into diverse data environments. DataCap's system architecture is designed to handle complex data workflows, providing a robust framework for data integration and visualization tasks. The software is suitable for users who need to consolidate data from multiple sources, transform it according to their needs, and visualize the results for better insights. It is particularly useful in big data contexts where managing and analyzing large volumes of data from heterogeneous sources is critical. The project is open source and actively maintained, with features such as Docker support for easy deployment. It also includes user authentication with predefined usernames and passwords for initial access. The community around DataCap is active, with ongoing contributions and updates, making it a reliable choice for organizations looking to enhance their data processing capabilities. Overall, DataCap stands out as a powerful tool for data professionals seeking an integrated solution for data transformation, integration, and visualization, with strong support for ClickHouse and other major databases.

https://github.com/devlive-community/datacap

big-dataclickhousedata-chartsdata-integrationdata-monitoringdata-sourcesdata-transformationdata-visualizationdatabasedatacapdb2dockerdremiodruidelasticsearchh2hiveignitekylinkyuubimonetdbmongodbmysqlnosql-databaseoraclephoenixpostgresqlprestoredisrelational-databasesql-serversqlservertrino

Also in ETL and Data Processing

PeerDB-io/peerdb

PeerDB is a high-performance, PostgreSQL-optimized ETL tool that enables fast, reliable, and cost-effective streaming of data from Postgres to data warehouses, queues, and storage engines, with native integration in ClickHouse Cloud.

ytsaurus/ytsaurus

YTsaurus is a scalable, fault-tolerant open-source big data platform featuring MapReduce, SQL engine, NoSQL store, and integration with ClickHouse for fast analytics.

FrigadeHQ/trench

Trench is an open-source, production-ready analytics infrastructure built on ClickHouse and Kafka for scalable, real-time event tracking and analytics with GDPR compliance.

apache/gluten

Gluten is a middle layer that offloads JVM-based SQL engines' execution, such as Spark SQL, to high-performance native engines like ClickHouse and Velox, leveraging vectorized processing for accele...

wgzhao/Addax

Addax is a versatile and extensible open-source ETL tool that supports seamless data transfer between over 20 SQL and NoSQL data sources, including ClickHouse, with easy configuration and deployment options.

zerodha/dungbeetle

DungBeetle is a distributed job server for asynchronously queuing and executing heavy SQL read jobs on MySQL, PostgreSQL, and ClickHouse databases, designed to offload report generation and improve application performance.

ClickHouse/ClickBench

ClickBench is a comprehensive and reproducible benchmark designed to evaluate the performance of analytical databases, including ClickHouse, using realistic workloads derived from real-world web analytics data.

Canner/vulcan-sql

VulcanSQL is an Analytical Data API Framework that simplifies and accelerates the creation of secure, scalable RESTful APIs from databases and data warehouses for AI agents and data applications.