ytsaurus/ytsaurus
YTsaurus is a scalable, fault-tolerant open-source big data platform featuring MapReduce, SQL engine, NoSQL store, and integration with ClickHouse for fast analytics.
Awesome ClickHouse › ETL and Data Processing
PeerDB is a fast, simple, and cost-effective ETL/ELT tool designed specifically for replicating and streaming data from PostgreSQL to various data warehouses, queues, and storage engines. It is optimized for PostgreSQL users who need to move large volumes of data efficiently and reliably. PeerDB supports multiple streaming modes including log-based Change Data Capture (CDC), cursor-based streaming, and XMIN-based streaming, enabling real-time data synchronization with high throughput and low latency. The tool is engineered to be 10 times faster than existing solutions by leveraging native PostgreSQL features such as comprehensive data type support (including jsonb, arrays, and geospatial data), efficient handling of large TOAST columns, and schema change management. PeerDB's architecture allows parallelization of initial data loads for large tables, significantly reducing synchronization times from days to minutes. It also incorporates fault tolerance mechanisms such as state management, automatic retries, idempotency handling, and configurable batching and parallelism to prevent crashes and out-of-memory errors. The project provides a PostgreSQL-compatible SQL interface for ETL operations, allowing users to utilize familiar tools and ecosystems like pgAdmin, psql, BI tools (Grafana, Tableau), migration/versioning tools (Flyway), and programming languages (Python, Go, Node.js) for development and scheduling. PeerDB is integrated natively with ClickHouse Cloud, offering a seamless Postgres CDC connector experience. It uses MinIO for staging data files internally and requires network access configuration when ClickHouse runs outside Docker environments. The project is open source under the Elastic License 2.0 and maintains active community support through Slack and comprehensive documentation. PeerDB aims to simplify and optimize data replication workflows for PostgreSQL-centric data stacks, making it a valuable tool for organizations relying heavily on PostgreSQL for their data infrastructure.
https://github.com/PeerDB-io/peerdb
YTsaurus is a scalable, fault-tolerant open-source big data platform featuring MapReduce, SQL engine, NoSQL store, and integration with ClickHouse for fast analytics.
Trench is an open-source, production-ready analytics infrastructure built on ClickHouse and Kafka for scalable, real-time event tracking and analytics with GDPR compliance.
Gluten is a middle layer that offloads JVM-based SQL engines' execution, such as Spark SQL, to high-performance native engines like ClickHouse and Velox, leveraging vectorized processing for accele...
Addax is a versatile and extensible open-source ETL tool that supports seamless data transfer between over 20 SQL and NoSQL data sources, including ClickHouse, with easy configuration and deployment options.
DungBeetle is a distributed job server for asynchronously queuing and executing heavy SQL read jobs on MySQL, PostgreSQL, and ClickHouse databases, designed to offload report generation and improve application performance.
ClickBench is a comprehensive and reproducible benchmark designed to evaluate the performance of analytical databases, including ClickHouse, using realistic workloads derived from real-world web analytics data.
DataCap is an integrated software platform for data transformation, integration, and visualization, supporting a wide range of data sources including ClickHouse and other major databases.
VulcanSQL is an Analytical Data API Framework that simplifies and accelerates the creation of secure, scalable RESTful APIs from databases and data warehouses for AI agents and data applications.