Apache Hudi (pronounced Hoodie) stands for Hadoop Upserts Deletes and Incrementals. Hudi manages the storage of large analytical datasets on DFS (Cloud stores, HDFS or any Hadoop FileSystem compatible storage). As an organization, Hudi can help you build an efficient data lake, solving some of the most complex, low-level storage management problems, while putting data into hands of your data analysts, engineers and scientists much quicker.
<BR><BR>
Features: <BR>
<ul>
<li>Upsert support with fast, pluggable indexing</li>
<li>Atomically publish data with rollback support</li>
<li>Snapshot isolation between writer &amp; queries</li>
<li>Savepoints for data recovery</li>
<li>Manages file sizes, layout using statistics</li>
<li>Async compaction of row &amp; columnar data</li>
<li>Timeline metadata to track lineage</li>
<li>Optimize data lake layout with clustering</li>
</ul>

Apache Hudi (pronounced Hoodie) stands for Hadoop Upserts Deletes and Incrementals. Hudi manages the storage of large analytical datasets on DFS (Cloud stores, HDFS or any Hadoop FileSystem compatible storage). As an organization, Hudi can help you build an efficient data lake, solving some of the most complex, low-level storage management problems, while putting data into hands of your data analysts, engineers and scientists much quicker.

Apache Hudi - Streaming Data Lake Platform

Discover open source projects across all platforms

Projects

Apache Hudi - Streaming Data Lake Platform

TechStack

Tagcloud

License

Suggested keywords:

Projects

Apache Hudi - Streaming Data Lake Platform

TechStack

Tagcloud

License