Awesome ClickHouseETL and Data Processing

FrigadeHQ/trench

⭐ 1664 TypeScript added to this list on 2024-12-11 repository created 2024-09-30

Trench is an open-source analytics infrastructure designed to provide a scalable, real-time event tracking system. Built on top of Apache Kafka and ClickHouse, Trench enables the processing and analysis of large volumes of event data efficiently. It is packaged as a single production-ready Docker image, making deployment straightforward and accessible. The system supports tracking events and page views, allowing users to build various analytics products such as product analytics dashboards, observability platforms, and LLM retrieval-augmented generation (RAG) systems. Trench is compliant with privacy regulations including GDPR and PECR, ensuring that users have full control over their data, including access, rectification, and deletion. It also supports the Segment API, providing compatibility with existing event tracking standards. The platform is capable of processing thousands of events per second on a single node and offers real-time querying capabilities, which is critical for timely insights and decision-making. Deployment options include a self-hosted version, which users can run on their own infrastructure using Docker and Docker Compose, and a fully-managed cloud solution that offers serverless operation, autoscaling, and high availability with 99.99% SLAs. The self-hosted version includes a local ClickHouse and Kafka instance, and users can interact with the system via RESTful API endpoints for sending events and querying data. Trench also supports webhooks to connect data to other destinations. The project is open-source under the MIT License and has an active community supported through Slack. It provides comprehensive documentation and demo resources to help users get started quickly. Trench was developed by the team at Frigade to meet the needs of scalable, real-time event tracking and analytics.

https://github.com/FrigadeHQ/trench

apache-kafkaautoscalingclickhouseclouddockerevent-trackingfrigadegdpr-compliantllm-ragmit-licenseobservabilityopen-sourcepecr-compliantproduct-analyticsreal-time-analyticsrest-apisegment-apiself-hostedserverlesswebhooks

Also in ETL and Data Processing

PeerDB-io/peerdb

PeerDB is a high-performance, PostgreSQL-optimized ETL tool that enables fast, reliable, and cost-effective streaming of data from Postgres to data warehouses, queues, and storage engines, with native integration in ClickHouse Cloud.

ytsaurus/ytsaurus

YTsaurus is a scalable, fault-tolerant open-source big data platform featuring MapReduce, SQL engine, NoSQL store, and integration with ClickHouse for fast analytics.

apache/gluten

Gluten is a middle layer that offloads JVM-based SQL engines' execution, such as Spark SQL, to high-performance native engines like ClickHouse and Velox, leveraging vectorized processing for accele...

wgzhao/Addax

Addax is a versatile and extensible open-source ETL tool that supports seamless data transfer between over 20 SQL and NoSQL data sources, including ClickHouse, with easy configuration and deployment options.

zerodha/dungbeetle

DungBeetle is a distributed job server for asynchronously queuing and executing heavy SQL read jobs on MySQL, PostgreSQL, and ClickHouse databases, designed to offload report generation and improve application performance.

ClickHouse/ClickBench

ClickBench is a comprehensive and reproducible benchmark designed to evaluate the performance of analytical databases, including ClickHouse, using realistic workloads derived from real-world web analytics data.

devlive-community/datacap

DataCap is an integrated software platform for data transformation, integration, and visualization, supporting a wide range of data sources including ClickHouse and other major databases.

Canner/vulcan-sql

VulcanSQL is an Analytical Data API Framework that simplifies and accelerates the creation of secure, scalable RESTful APIs from databases and data warehouses for AI agents and data applications.