Semantic Coherence: Apache Kudu – Signal Evidence & AI Readability

Apache Kudu

(https://kudu.apache.org) 📸 Data Snapshot: May 27, 2026
Semantic Coherence — The Lens

Pull the main entities out of the H1, then check whether they actually recur through the body. A page that announces one thing and then talks about another drifts. Headings with no real sentences underneath read as pseudo-substance.

Semantic Coherence Homepage promise vs. Sub-page reality.
20 Impact Weight: 20 / 100
100% Reputation

There is zero semantic drift between the homepage and sub-pages. The homepage hero signal promising a distributed data storage engine is rigorously supported by the Overview and Docs pages, which provide deep architectural dives into columnar storage and Raft replication. The messaging is consistent for a developer and architect audience, maintaining a high-level technical tone without pivoting to generic business benefits.

Semantic Coherence is read from the heading hierarchy first: what each page announces in its H1 and headings, then whether the body actually delivers on it. Below is the structure the engine mapped, followed by the clean text to check for drift between promise and reality.

🏗️ Semantic Structure — heading hierarchy & page identity (the promise the page makes)
HOMEPAGE Apache Kudu – Fast Analytics on Fast Data (https://kudu.apache.org)
Title

Apache Kudu – Fast Analytics on Fast Data

Meta

A new open source Apache Hadoop ecosystem project, Apache Kudu completes Hadoop

H3 Streamlined Architecture
H3 Faster Analytics
H3 Open for Contributions
NAV_REPEATED Apache Kudu – Community (https://kudu.apache.org/community.html)
Title

Apache Kudu – Community

Meta

A new open source Apache Hadoop ecosystem project, Apache Kudu completes Hadoop

H3 Apache Kudu Mailing Lists and Chat Rooms
H3 Contributions
H3 Meetups, User Groups, and Conference Presentations
H3 Articles, Demos, and Reviews
H4 Participate in the community.
H4 Talk about how you use Kudu.
H4 File bugs and enhancement requests.
H4 Write code and submit patches.
H4 Review patches and test new code.
H4 Write and review documentation or a blog post.
H4 Request and review examples.
NAV Apache Kudu – Overview (https://kudu.apache.org/overview.html)
Title

Apache Kudu – Overview

Meta

A new open source Apache Hadoop ecosystem project, Apache Kudu completes Hadoop

H2 Apache Kudu Overview
H2 Kudu Architecture
H3 Data Model
H3 Low-latency random access
H3 Apache Hadoop Ecosystem Integration
H3 Built by and for Operators
H3 Open Source
H3 Super-fast Columnar Storage
H3 Distribution and Fault Tolerance
H3 Designed for Next-Generation Hardware
NAV Apache Kudu – Introducing Apache Kudu (https://kudu.apache.org/docs/)
Title

Apache Kudu – Introducing Apache Kudu

Meta

A new open source Apache Hadoop ecosystem project, Apache Kudu completes Hadoop

H1 Introducing Apache Kudu
H2 Kudu-Impala Integration Features
H2 Concepts and Terms
H2 Architectural Overview
H2 Example Use Cases
H2 Next Steps
📝 The Narrative — clean text per page (homepage promise vs. sub-page reality)
HOMEPAGE (https://kudu.apache.org) Apache Kudu – Fast Analytics on Fast Data
[IMG: Apache Kudu]
Apache Kudu is an open source distributed data storage engine that makes fast analytics on fast and changing data easy.
Quickstart
Installation
Releases

[H3] Streamlined Architecture
Kudu provides a combination of fast inserts/updates and efficient columnar scans to enable multiple real-time analytic workloads across a single storage layer. Kudu gives architects the flexibility to address a wider variety of use cases without exotic workarounds and no required external service dependencies.
Learn more »

[H3] Faster Analytics
Kudu is specifically designed for use cases that require fast analytics on fast (rapidly changing) data. Engineered to take advantage of next-generation hardware and in-memory processing, Kudu lowers query latency significantly for engines like Apache Impala, Apache NiFi, Apache Spark, Apache Flink, and more.
Learn more »

[H3] Open for Contributions
Founded by long-time contributors to the Apache big data ecosystem, Apache Kudu is a top-level Apache Software Foundation project released under the Apache 2 license and values community participation as an important ingredient in its long-term success. We appreciate all community contributions to date, and are looking forward to seeing more!
Learn more »
1274 chars
SUB-PAGE (https://kudu.apache.org/community.html) Apache Kudu – Community
[H3] Apache Kudu Mailing Lists and Chat Rooms
Get help using Kudu or contribute to the project on our mailing lists or our chat room:
user@kudu.apache.org
(subscribe)
(unsubscribe)
for usage questions, help, and announcements.
Kudu Slack channel -
where many Kudu developers and users hang out to answer questions and chat.
Developer mailing lists
dev@kudu.apache.org
(subscribe)
(unsubscribe)
for people who want to contribute code to Kudu.
builds@kudu.apache.org
(subscribe)
(unsubscribe)
for discussions and notifications surrounding build infrastructure.
issues@kudu.apache.org
(subscribe)
(unsubscribe)
receives an email notification for all ticket updates made in the Kudu JIRA issue tracker.
reviews@kudu.apache.org
(subscribe)
(unsubscribe)
receives an email notification for all code review requests and responses on the
Kudu Gerrit.
commits@kudu.apache.org
(subscribe)
(unsubscribe)
receives an email notification of all code changes to the
Kudu Git repository.
Other developer resources
GitHub
Gerrit Code Review
JIRA Issue Tracker
Social Media
Twitter
Reddit
Project information
Apache Kudu Committers list
Apache Kudu Ecosystem
Security
Sponsorship
Thanks
License
[H3] Contributions
There are lots of ways to get involved with the Kudu project. Some of them are
listed below. You don’t have to be a developer; there are lots of valuable and
important ways to get involved that suit any skill set and level.
If you want to do something not listed here, or you see a gap that needs to be
filled, let us know.
[H4] Participate in the community.
Community is the core of any open source project, and Kudu is no exception.
Participate in the mailing lists, requests for comment, chat sessions, and bug
reports.
[H4] Talk about how you use Kudu.
Let us know what you think of Kudu and how you are using it. Send links to
blogs or presentations you’ve given to the kudu user mailing
list so that we can feature them.
[H4] File bugs and enhancement requests.
If you see problems in Kudu or if a missing feature would make Kudu more useful
to you, let us know by filing a bug or request for enhancement on the Kudu
JIRA issue tracker. The more
information you can provide about how to reproduce an issue or how you’d like a
new feature to work, the better.
[H4] Write code and submit patches.
You can submit patches to the core Kudu project or extend your existing
codebase and APIs to work with Kudu. The Kudu project uses
Gerrit for code
reviews. Please read the details of how to submit
patches and what
the project coding guidelines are before
your submit your patch, so that your contribution will be easy for others to
review and integrate.
[H4] Review patches and test new code.
In order for patches to be integrated into Kudu as quickly as possible, they
must be reviewed and tested. The more eyes, the better. Even if you are not a
committer your review input is extremely valuable. Keep an eye on the Kudu
gerrit instance
for patches that need review or testing.
[H4] Write and review documentation or a blog post.
Making good documentation is critical to making great, usable software. If you
see gaps in the documentation, please submit suggestions or corrections to the
mailing list or submit documentation patches through Gerrit. You can also
correct or improve error messages, log messages, or API docs.
If you’d like to translate the Kudu documentation into a different language or
you’d like to help in some other way, please let us know.
It’s best to review the documentation guidelines
before you get started.
[H4] Request and review examples.
The examples directory
includes working code examples. As more examples are requested and added, they
will need review and clean-up. This is another way you can get involved.
[H3] Meetups, User Groups, and Conference Presentations
If you’re interested in hosting or presenting a Kudu-related talk or meetup in
your city, get in touch by sending email to the user mailing list at
user@kudu.apache.org
so that we can feature them.
Presentations about Kudu are planned or have taken place at the following events:
Thu, Oct 27, 2016. Spark Summit EU. Brussels, Belgium.
Apache Kudu and Spark SQL. Presented by Mike Percy.
Tue, Sep 6, 2016. Boston Cloudera User Group. Boston, MA, USA.
Apache Kudu 0.10 and Spark SQL. Presented by William Berkeley.
Thu, Aug 18, 2016. Boulder/Denver Big Data Meetup. Broomfield, CO, USA.
Apache Kudu: New Apache Hadoop Storage for Fast Analytics on Fast Data. Presented by Mike Percy.
Wed, Aug 17, 2016. Denver Cloudera User Group. Greenwood Village, CO, USA.
Apache Kudu: New Apache Hadoop Storage for Fast Analytics on Fast Data. Presented by Mike Percy.
Thu, July 21, 2016. Silicon Valley Big Data Meetup. Palo Alto, CA, USA.
Apache Kudu (incubating): New Apache Hadoop Storage for Fast Analytics on Fast Data. Presented by Mike Percy.
Wed, May 25, 2016. Dallas/Fort Worth Cloudera User Group. Dallas, TX, USA.
Apache Kudu: New Apache Hadoop Storage for Fast Analytics on Fast Data. Presented by Ryan Bosshart.
Thu, May 12, 2016. Vancouver Spark Meetup. Vancouver, BC, Canada.
Kudu and Spark for fast analytics on streaming data. Presented by Mike Percy and Dan Burkert.
Tue, May 10, 2016. Apache: Big Data 2016. Vancouver, BC, Canada.
Using Kafka and Kudu for fast, low-latency SQL analytics on streaming data. Presented by Mike Percy and Ashish Singh.
Mon, May 9, 2016. Apache: Big Data 2016. Vancouver, BC, Canada.
Introduction to Apache Kudu (incubating) for timeseries storage. Presented by Dan Burkert.
Wed, May 4, 2016. Cloudera Tech Meetup. Budapest, Hungary.
Apache Kudu (incubating): New Apache Hadoop Storage for Fast Analytics on Fast Data. Presented by Adar Dembo.
Fri, Apr 8, 2016. DataEngConf. San Francisco, CA, USA.
Resolving Transactional Access/Analytic Performance Trade-offs in Apache Hadoop with Apache Kudu.
Thu, Mar 31, 2016. Strata Hadoop World. San Jose, CA, USA.
Fast data made easy with Apache Kafka and Apache Kudu (incubating). Presented by Ted Malaska and Jeff Holoman.
Sat, Mar 19, 2016. China Hadoop Summit 2016. Beijing, China.
使用Kudu构建统一实时数据分析服务的方法 (Using Kudu to build unified real-time data analysis services). Presented by Binglin Chang.
Thu, Mar 17, 2016. Big Data Boston. Boston, MA, USA.
St. Patty’s Day meet-up on an Introduction to Apache Kudu. Presented by Todd Lipcon.
Wed, Mar 16, 2016. Cloudera Technology Day. Washington DC, USA.
Apache Kudu (Incubating): New Hadoop Storage for Fast Analytics on Fast Data. Presented by Todd Lipcon.
Tue, Mar 1, 2016. Rust Detroit. Detroit, MI, USA.
Hadoop Next Gen: Using Kudu & Mozilla Rust to Crunch Big Data!. Presented by Dan Burkert.
Wed, Feb 24, 2016. Seattle Scalability Meetup. Seattle, WA, USA.
Resolving Transactional Access/Analytic Performance Trade-offs in Hadoop with Kudu. Presented by Dan Burkert.
Tue, Feb 23, 2016. SF Data Engineering Meetup. San Francisco, CA, USA.
Intro to Apache Kudu. Presented by Asim Jalis. (slides)
Thu, Feb 18, 2016. DataKRK Meetup. Krakow, Poland.
Are you KUDUing me?. Presented by Przemek Maciołek. (slides)
Wed, Feb 17, 2016. Bay Area Hadoop User Group. Sunnyvale, CA, USA.
Apache Kudu (incubating): New Apache Hadoop Storage for Fast Analytics on Fast Data. Presented by David Alves. (video)
Wed, Feb 10, 2016. San Francisco Python Meetup Group. San Francisco, CA, USA.
Using Python at Scale for Data Science (Python + Kudu + Ibis). Presented by Wes McKinney.
Mon, Feb 8, 2016. Hadoop / Spark Conference Japan 2016. Tokyo, Japan.
KuduによるHadoopのトランザクションアクセスと分析パフォーマンスのトレードオフ解消. Presented by Todd Lipcon.
Wed, Jan 27, 2015. Big Data Application Meetup. Palo Alto, CA, USA.
Simplifying big data analytics with Apache Kudu. Presented by Mike Percy. (video)
Tue, Dec 15, 2015. San Francisco Spark Hackers. San Francisco, CA, USA.
Faster than Parquet! A deep dive into Kudu. Presented by Jean-Daniel Cryans. (video)
Thu, Dec 10, 2015. Big Data Technology Conference Beijing. Beijing, China.
Kudu: Fast analytics on fast data. Presented by Todd Lipcon (Cloudera).
Wed, Dec 09, 2015. The Hive Big Data Think Tank. Palo Alto, CA, USA.
Kudu: New Apache Hadoop Storage for Fast Analytics on Fast Data. Presented by Mike Percy (Cloudera). (video)
Tue, Dec 08, 2015. Korea Big Data Think Tank. Seoul, Korea (South).
Kudu: New Apache Hadoop Storage for Fast Analytics on Fast Data. Presented by Todd Lipcon (Cloudera).
Sun, Dec 06, 2015. Shanghai Big Data Streaming Meetup. Shanghai, China.
Kudu: Fast analytics on fast data. Presented by Todd Lipcon (Cloudera).
Wed, Dec 02, 2015. Kudu office hours at Strata Singapore. Singapore.
Office hour with Todd Lipcon (Cloudera).
Wed, Dec 02, 2015. Strata Singapore. Singapore.
Hadoop’s storage gap: Resolving transactional access/analytic performance trade-offs with Kudu.
Presented by Todd Lipcon (Cloudera).
Thu, Nov 05, 2015. Washington DC Area Apache Spark Interactive. Washington, DC, USA.
A Spark Auto Scaling Kudu Sneak Peek in 3s.
Presented by Jean-Daniel Cryans (Cloudera).
Thu, Oct 22, 2015. SF Spark and Friends. San Francisco, CA, USA.
Kudu: Data Store for the New Era, with Kafka+Spark+Kudu Demo.
Presented by Jean-Daniel Cryans (Cloudera).
Tue, Oct 06, 2015. San Francisco Hadoop User Group. San Francisco, CA, USA.
Resolving Transactional Access/Analytic Performance Trade-offs in Apache Hadoop.
Presented by Todd Lipcon (Cloudera).
Thu, Oct 01, 2015. Strata New York. New York, NY, USA.
Ask Me Anything Panel with the Kudu Development Team.
Presented by members of the Kudu development team.
Wed, Sep 30, 2015. Strata New York. New York, NY, USA.
Resolving Transactional Access/Analytic Performance Trade-offs in Apache Hadoop.
Presented by Todd Lipcon (Cloudera) and Binglin Chang (Xiaomi).
Tue, Sep 29, 2015. NYC Hadoop User Group. New York, NY, USA.
Resolving Transactional Access/Analytic Performance Trade-offs in Apache Hadoop.
Presented by Todd Lipcon (Cloudera).
[H3] Articles, Demos, and Reviews
The Kudu community does not yet have a dedicated blog, but if you are
interested in promoting a Kudu-related use case, we can help spread the word.
Send email to the user mailing list at
user@kudu.apache.org
with your content and we’ll help drive traffic.
Curt Monash from DBMS2 has written a three-part series about Kudu:
an introduction to Kudu,
a technical deep dive,
and his analysis on
the potential significance of Kudu.
Zoomdata has created a video demo of
Zoomdata on top of Kudu
demonstrating real-time and point-in-time analytic queries on Kudu while
simultaneously running a streaming ingest workload.
10573 chars
SUB-PAGE (https://kudu.apache.org/overview.html) Apache Kudu – Overview
[H3] Data Model

A Kudu cluster stores tables that look just like tables you're used to from relational (SQL) databases.
A table can be as simple as an binary key and value, or as complex
as a few hundred different strongly-typed attributes.

Just like SQL, every table has a PRIMARY KEY made up of one or more columns.
This might be a single column like a unique user identifier, or a compound key such as a
(host, metric, timestamp) tuple for a machine time series database. Rows can be efficiently
read, updated, or deleted by their primary key.

Kudu's simple data model makes it breeze to port legacy applications or build new ones:
no need to worry about how to encode your data into binary blobs or make sense of a
huge database full of hard-to-interpret JSON. Tables are self-describing, so you can
use standard tools like SQL engines or Spark to analyze your data.

Learn more about schema design with Kudu

[H3] Low-latency random access

Unlike other storage for big data analytics, Kudu isn't just a file format. It's a live storage
system which supports low-latency millisecond-scale access to individual rows. For
"NoSQL"-style access, you can choose between Java, C++, or Python APIs. And of course these
random access APIs can be used in conjunction with batch access for machine learning or analytics.

Kudu's APIs are designed to be easy to use. The data model is fully typed, so you don't
need to worry about binary encodings or exotic serialization. You can just store primitive
types, like when you use JDBC or ODBC.

Kudu isn't designed to be an OLTP system, but if you have some subset of data which fits
in memory, it offers competitive random access performance. We've measured 99th percentile
latencies of 6ms or below using YCSB with a uniform random access workload over a billion
rows. Being able to run low-latency online workloads on the same storage as back-end
data analytics can dramatically simplify application architecture.

View the Java API docs
View the C++ API docs
Learn more about developing applications with Kudu

[H3] Apache Hadoop Ecosystem Integration

Kudu was designed to fit in with the Hadoop ecosystem, and integrating it with other
data processing frameworks is simple. You can stream data in from live real-time data sources
using the Java client, and then process it immediately upon arrival using Spark, Impala,
or MapReduce. You can even transparently join Kudu tables with data stored in other Hadoop
storage such as HDFS or HBase.

Kudu is a good citizen on a Hadoop cluster: it can easily share data
disks with HDFS DataNodes, and can operate in a RAM footprint as small as 1 GB for
light workloads.

Learn more about integration with Impala
View an example of a MapReduce job on Kudu

[H3] Built by and for Operators

Kudu was built by a group of engineers who have spent many late nights providing
on-call production support for critical Hadoop clusters across hundreds of
enterprise use cases. We know how frustrating it is to debug software
without good metrics, tracing, or administrative tools.

Ever since its first beta release, Kudu has included advanced in-process tracing capabilities,
extensive metrics support, and even watchdog threads which check for latency
outliers and dump "smoking gun" stack traces to get to the root of the problem
quickly.

Learn more about administering Kudu
Learn more about Kudu's tracing capabilities

[H3] Open Source

Kudu is Open Source software, licensed under the Apache 2.0 license and
governed under the aegis of the Apache Software Foundation. We believe
that Kudu's long-term success depends on building a vibrant community of
developers and users from diverse organizations and backgrounds.

Learn more about how to contribute
View the Kudu github repository

[H3] Super-fast Columnar Storage

Like most modern analytic data stores, Kudu internally organizes its data by column rather than
row. Columnar storage allows efficient encoding and compression. For example, a string field with
only a few unique values can use only a few bits per row of storage. With techniques such as
run-length encoding, differential encoding, and vectorized bit-packing, Kudu is as fast at
reading the data as it is space-efficient at storing it.

Columnar storage also dramatically reduces the amount of data IO required to service analytic
queries. Using techniques such as lazy data materialization and predicate pushdown, Kudu can perform
drill-down and needle-in-a-haystack queries over billions of rows and terabytes of data in seconds.

Read the Kudu paper for more details and a performance evaluation

[H3] Distribution and Fault Tolerance

In order to scale out to large datasets and large clusters, Kudu splits tables
into smaller units called tablets. This splitting can be configured
on a per-table basis to be based on hashing, range partitioning, or a combination
thereof. This allows the operator to easily trade off between parallelism for
analytic workloads and high concurrency for more online ones.

In order to keep your data safe and available at all times, Kudu uses the
Raft consensus algorithm to replicate
all operations for a given tablet. Raft, like Paxos, ensures that every
write is persisted by at least two nodes before responding to
the client request, ensuring that no data is ever lost due to a
machine failure. When machines do fail, replicas reconfigure
themselves within a few seconds to maintain extremely high system
availability.

The use of majority consensus provides very low tail latencies
even when some nodes may be stressed by concurrent workloads such as
Spark jobs or heavy Impala queries. But unlike eventually
consistent systems, Raft consensus ensures that all replicas will
come to agreement around the state of the data, and by using a
combination of logical and physical clocks, Kudu can offer strict
snapshot consistency to clients that demand it.

Learn more about Raft Consensus
Read the Kudu paper for more details on its architecture

[H3] Designed for Next-Generation Hardware

The Kudu team has worked closely with engineers at Intel to harness the power
of the next generation of hardware technologies.
Kudu's storage is designed to take advantage of the IO
characteristics of solid state drives, and it includes an
experimental cache implementation based on the libpmem
library which can store data in persistent memory.

Kudu is implemented in C++, so it can scale easily to large amounts
of memory per node. And because key storage data structures are designed to
be highly concurrent, it can scale easily to tens of cores. With an
in-memory columnar execution path, Kudu achieves good instruction-level
parallelism using SIMD operations from the SSE4 and AVX instruction sets.
6899 chars
SUB-PAGE (https://kudu.apache.org/docs/) Apache Kudu – Introducing Apache Kudu
[H1] Introducing Apache Kudu

Kudu is a distributed columnar storage engine optimized for OLAP workloads.
Kudu runs on commodity hardware, is horizontally scalable, and supports highly
available operation.
Kudu’s design sets it apart. Some of Kudu’s benefits include:
Fast processing of OLAP workloads.
Strong but flexible consistency model, allowing you to choose consistency
requirements on a per-request basis, including the option for
strict-serializable consistency.
Structured data model.
Strong performance for running sequential and random workloads simultaneously.
Tight integration with Apache Impala, making it a good, mutable alternative to
using HDFS with Apache Parquet.
Integration with Apache NiFi and Apache Spark.
Integration with Hive Metastore (HMS) and Apache Ranger to provide
fine-grain authorization and access control.
Authenticated and encrypted RPC communication.
High availability: Tablet Servers and Masters use the Raft Consensus Algorithm, which ensures
that as long as more than half the total number of tablet replicas is
available, the tablet is available for reads and writes. For instance,
if 2 out of 3 replicas (or 3 out of 5 replicas, etc.) are available,
the tablet is available. Reads can be serviced by read-only follower tablet
replicas, even in the event of a leader replica’s failure.
Automatic fault detection and self-healing: to keep data highly available,
the system detects failed tablet replicas and re-replicates data from
available ones, so failed replicas are automatically replaced when enough
Tablet Servers are available in the cluster.
Location awareness (a.k.a. rack awareness) to keep the system available
in case of correlated failures and allowing Kudu clusters to span over
multiple availability zones.
Logical backup (full and incremental) and restore.
Multi-row transactions (only for INSERT/INSERT_IGNORE operations as of
Kudu 1.15 release).
Easy to administer and manage.
By combining all of these properties, Kudu targets support for families of
applications that are difficult or impossible to implement using Hadoop storage
technologies, while it is compatible with most of the data processing
frameworks in the Hadoop ecosystem.
A few examples of applications for which Kudu is a great solution are:
Reporting applications where newly-arrived data needs to be immediately available for end users
Time-series applications that must simultaneously support:
queries across large amounts of historic data
granular queries about an individual entity that must return very quickly
Applications that use predictive models to make real-time decisions with periodic
refreshes of the predictive model based on all historic data
For more information about these and other scenarios, see Example Use Cases.
[H2] Kudu-Impala Integration Features
CREATE/ALTER/DROP TABLE
Impala supports creating, altering, and dropping tables using Kudu as the persistence layer.
The tables follow the same internal / external approach as other tables in Impala,
allowing for flexible data ingestion and querying.
INSERT
Data can be inserted into Kudu tables in Impala using the same syntax as
any other Impala table like those using HDFS or HBase for persistence.
UPDATE / DELETE
Impala supports the UPDATE and DELETE SQL commands to modify existing data in
a Kudu table row-by-row or as a batch. The syntax of the SQL commands is chosen
to be as compatible as possible with existing standards. In addition to simple DELETE
or UPDATE commands, you can specify complex joins with a FROM clause in a subquery.
Flexible Partitioning
Similar to partitioning of tables in Hive, Kudu allows you to dynamically
pre-split tables by hash or range into a predefined number of tablets, in order
to distribute writes and queries evenly across your cluster. You can partition by
any number of primary key columns, by any number of hashes, and an optional list of
split rows. See Schema Design.
Parallel Scan
To achieve the highest possible performance on modern hardware, the Kudu client
used by Impala parallelizes scans across multiple tablets.
High-efficiency queries
Where possible, Impala pushes down predicate evaluation to Kudu, so that predicates
are evaluated as close as possible to the data. Query performance is comparable
to Parquet in many workloads.
For more details regarding querying data stored in Kudu using Impala, please
refer to the Impala documentation.
[H2] Concepts and Terms
Columnar Data Store
Kudu is a columnar data store. A columnar data store stores data in strongly-typed
columns. With a proper design, it is superior for analytical or data warehousing
workloads for several reasons.
Read Efficiency
For analytical queries, you can read a single column, or a portion
of that column, while ignoring other columns. This means you can fulfill your query
while reading a minimal number of blocks on disk. With a row-based store, you need
to read the entire row, even if you only return values from a few columns.
Data Compression
Because a given column contains only one type of data,
pattern-based compression can be orders of magnitude more efficient than
compressing mixed data types, which are used in row-based solutions. Combined
with the efficiencies of reading data from columns, compression allows you to
fulfill your query while reading even fewer blocks from disk. See
Data Compression
Table
A table is where your data is stored in Kudu. A table has a schema and
a totally ordered primary key. A table is split into segments called tablets.
Tablet
A tablet is a contiguous segment of a table, similar to a partition in
other data storage engines or relational databases. A given tablet is
replicated on multiple tablet servers, and at any given point in time,
one of these replicas is considered the leader tablet. Any replica can service
reads, and writes require consensus among the set of tablet servers serving the tablet.
Tablet Server
A tablet server stores and serves tablets to clients. For a
given tablet, one tablet server acts as a leader, and the others act as
follower replicas of that tablet. Only leaders service write requests, while
leaders or followers each service read requests. Leaders are elected using
Raft Consensus Algorithm. One tablet server can serve multiple tablets, and one tablet can be served
by multiple tablet servers.
Master
The master keeps track of all the tablets, tablet servers, the
Catalog Table, and other metadata related to the cluster. At a given point
in time, there can only be one acting master (the leader). If the current leader
disappears, a new master is elected using Raft Consensus Algorithm.
The master also coordinates metadata operations for clients. For example, when
creating a new table, the client internally sends the request to the master. The
master writes the metadata for the new table into the catalog table, and
coordinates the process of creating tablets on the tablet servers.
All the master’s data is stored in a tablet, which can be replicated to all the
other candidate masters.
Tablet servers heartbeat to the master at a set interval (the default is once
per second).
Raft Consensus Algorithm
Kudu uses the Raft consensus algorithm as
a means to guarantee fault-tolerance and consistency, both for regular tablets and for master
data. Through Raft, multiple replicas of a tablet elect a leader, which is responsible
for accepting and replicating writes to follower replicas. Once a write is persisted
in a majority of replicas it is acknowledged to the client. A given group of N replicas
(usually 3 or 5) is able to accept writes with at most (N - 1)/2 faulty replicas.
Catalog Table
The catalog table is the central location for
metadata of Kudu. It stores information about tables and tablets. The catalog
table may not be read or written directly. Instead, it is accessible
only via metadata operations exposed in the client API.
The catalog table stores two categories of metadata:
Tables
table schemas, locations, and states
Tablets
the list of existing tablets, which tablet servers have replicas of
each tablet, the tablet’s current state, and start and end keys.
Logical Replication
Kudu replicates operations, not on-disk data. This is referred to as logical replication,
as opposed to physical replication. This has several advantages:
Although inserts and updates do transmit data over the network, deletes do not need
to move any data. The delete operation is sent to each tablet server, which performs
the delete locally.
Physical operations, such as compaction, do not need to transmit the data over the
network in Kudu. This is different from storage systems that use HDFS, where
the blocks need to be transmitted over the network to fulfill the required number of
replicas.
Tablets do not need to perform compactions at the same time or on the same schedule,
or otherwise remain in sync on the physical storage layer. This decreases the chances
of all tablet servers experiencing high latency at the same time, due to compactions
or heavy write loads.
[H2] Architectural Overview
The following diagram shows a Kudu cluster with three masters and multiple tablet
servers, each serving multiple tablets. It illustrates how Raft consensus is used
to allow for both leaders and followers for both the masters and tablet servers. In
addition, a tablet server can be a leader for some tablets, and a follower for others.
Leaders are shown in gold, while followers are shown in blue.
[IMG: Kudu Architecture]
[H2] Example Use Cases
Streaming Input with Near Real Time Availability
A common challenge in data analysis is one where new data arrives rapidly and constantly,
and the same data needs to be available in near real time for reads, scans, and
updates. Kudu offers the powerful combination of fast inserts and updates with
efficient columnar scans to enable real-time analytics use cases on a single storage layer.
Time-series application with widely varying access patterns
A time-series schema is one in which data points are organized and keyed according
to the time at which they occurred. This can be useful for investigating the
performance of metrics over time or attempting to predict future behavior based
on past data. For instance, time-series customer data might be used both to store
purchase click-stream history and to predict future purchases, or for use by a
customer support representative. While these different types of analysis are occurring,
inserts and mutations may also be occurring individually and in bulk, and become available
immediately to read workloads. Kudu can handle all of these access patterns
simultaneously in a scalable and efficient manner.
Kudu is a good fit for time-series workloads for several reasons. With Kudu’s support for
hash-based partitioning, combined with its native support for compound row keys, it is
simple to set up a table spread across many servers without the risk of "hotspotting"
that is commonly observed when range partitioning is used. Kudu’s columnar storage engine
is also beneficial in this context, because many time-series workloads read only a few columns,
as opposed to the whole row.
In the past, you might have needed to use multiple data stores to handle different
data access patterns. This practice adds complexity to your application and operations,
and duplicates your data, doubling (or worse) the amount of storage
required. Kudu can handle all of these access patterns natively and efficiently,
without the need to off-load work to other data stores.
Predictive Modeling
Data scientists often develop predictive learning models from large sets of data. The
model and the data may need to be updated or modified often as the learning takes
place or as the situation being modeled changes. In addition, the scientist may want
to change one or more factors in the model to see what happens over time. Updating
a large set of data stored in files in HDFS is resource-intensive, as each file needs
to be completely rewritten. In Kudu, updates happen in near real time. The scientist
can tweak the value, re-run the query, and refresh the graph in seconds or minutes,
rather than hours or days. In addition, batch or incremental algorithms can be run
across the data at any time, with near-real-time results.
Combining Data In Kudu With Legacy Systems
Companies generate data from multiple sources and store it in a variety of systems
and formats. For instance, some of your data may be stored in Kudu, some in a traditional
RDBMS, and some in files in HDFS. You can access and query all of these sources and
formats using Impala, without the need to change your legacy systems.
[H2] Next Steps
Get Started With Kudu
Installing Kudu

Introducing Kudu
Kudu-Impala Integration Features
Concepts and Terms
Architectural Overview
Example Use Cases
Next Steps

Kudu Release Notes

Quickstart Guide

Installation Guide

Configuring Kudu

Using the Hive Metastore with Kudu

Using Impala with Kudu

Administering Kudu

Troubleshooting Kudu

Developing Applications with Kudu

Kudu Schema Design

Kudu Scaling Guide

Kudu Security

Kudu Transaction Semantics

Background Maintenance Tasks

Kudu Configuration Reference

Kudu Command Line Tools Reference

Kudu Metrics Reference

Known Issues and Limitations

Contributing to Kudu

Export Control Notice
13349 chars