Data Pipelines

ROS 2 Data Logging at Fleet Scale: rosbag2, MCAP and Pipelines That Survive Production

Bimanual robot setup with two arms recording manipulation data

Recording a rosbag on a single robot on your desk is easy. Recording the right data, on fifty robots in three countries, every day, without filling disks, saturating uplinks or losing the one bag that explains last Tuesday's collision, is a data engineering problem. This guide covers how to configure rosbag2 and MCAP for production, and how to build the pipeline that sits behind them.

Key takeaways

  • Use the MCAP storage plugin (the rosbag2 default since ROS 2 Iron) and prefer its built-in chunk compression over rosbag2 file-mode compression, so bags stay seekable.
  • Split bags by size or duration so that every file is a small, independently uploadable and independently useful unit.
  • Combine a lean always-on recording profile with event-triggered snapshots from an in-memory circular buffer for high-bandwidth sensors.
  • Treat upload as a scheduling problem: bandwidth-aware, resumable, and prioritized by incident severity rather than by file age.
  • Index bags by robot, time, topic and event at ingest time, and extract to columnar formats like Parquet so analytics never has to open a raw bag.

Why ROS 2 data logging breaks at fleet scale

Most teams start ROS 2 data logging the same way: an engineer runs ros2 bag record -a, copies the result off the robot with scp, and opens it in a visualizer. That workflow is fine for a prototype. It fails at fleet scale for predictable reasons:

  • Volume explodes. A mobile robot with four compressed cameras, a 3D lidar and the usual odometry, TF and diagnostics can, for example, produce tens of gigabytes per operating hour. Across a fleet, recording everything stops being an option.
  • Disks fill silently. A recorder that runs until a 256 GB embedded disk is full can take down other processes, including the ones writing the logs you need for the post-mortem.
  • Uplinks are unreliable. Warehouse Wi-Fi, field LTE and strict customer networks interrupt uploads. Without resumable transfers you resend the same gigabytes repeatedly.
  • Nobody can find anything. Ten thousand bags named by timestamp in a bucket are not a dataset.
  • Recording affects the robot. Wrong QoS or CPU and I/O contention with perception nodes can drop messages or add control-loop latency.

A robust robot data pipeline therefore has two halves: rosbag2 configured to record the right topics, in the right format, in bounded space; and a fleet data management layer that moves, indexes, tiers and extracts that data.

rosbag2 storage plugins: sqlite3 vs MCAP

rosbag2 separates the recorder from the on-disk format through storage plugins. Two plugins matter in practice.

sqlite3 was the original default: each message is a row in a SQLite database. It is robust, but it carries per-message overhead, is not designed for streaming large binary payloads, and does not store the message definitions needed to decode it outside ROS.

MCAP is an open, self-describing container for timestamped, multi-channel messages, and the default rosbag2 storage since ROS 2 Iron. Its properties fit robotics well:

  • Self-describing. Message schemas are embedded, so a bag decodes years later without the original workspace.
  • Chunked and indexed. A summary section at the end of the file enables fast seeking and per-topic reads without a full scan.
  • Append-friendly. Built for streaming writes; a file truncated by power loss is usually recoverable up to the last complete chunk.
  • Language-neutral. Libraries exist for Python, C++, Go, Rust and TypeScript, so cloud extractors need no ROS installation.

The specification is documented at mcap.dev. On Humble, install the rosbag2_storage_mcap package; on any distribution, pass --storage mcap explicitly in launch files to avoid surprises.

Configuring rosbag2 for production: compression, splitting and QoS

The defaults of ros2 bag record are designed for interactive use. For production, three settings matter most.

Compression: zstd, file mode vs message mode vs MCAP chunks

rosbag2 has its own zstd compression layer with two modes. File mode compresses each completed file as a whole: good ratio, but the result must be decompressed before you can seek inside it. Message mode compresses each message individually: seekable, but small messages compress poorly.

With MCAP there is a better option: the storage plugin compresses whole chunks while keeping the index readable. Select it with a storage preset profile such as zstd_fast (lower CPU) or zstd_small (better ratio), or fastwrite to skip chunk compression when CPU is the bottleneck. Reserve rosbag2 file mode for bags going straight to cold storage.

Be realistic: structured topics such as odometry, TF and joint states typically compress 3–10x, while JPEG or H.264 payloads barely shrink. If you record raw sensor_msgs/Image, the biggest win is recording a compressed transport instead.

Splitting bags by size and duration

Never record one giant file per session. Use --max-bag-size (bytes) or --max-bag-duration (seconds) so the recorder rolls over regularly. Small files are easier to upload, retry and delete, and a corrupted file costs minutes of data rather than hours. Files of roughly 500 MB to 2 GB, or 1–5 minutes, are a good balance.

QoS overrides

The recorder adapts to publisher QoS, but be explicit for latched topics such as /tf_static (which need transient_local durability to capture messages published before recording started) and for high-rate sensors where best-effort avoids back-pressure:

# qos_overrides.yaml
/tf_static:
  history: keep_all
  reliability: reliable
  durability: transient_local
/robot_description:
  history: keep_last
  depth: 1
  reliability: reliable
  durability: transient_local
/lidar/points:
  history: keep_last
  depth: 10
  reliability: best_effort
  durability: volatile

Putting it together, a production recording command looks like this:

ros2 bag record \
  --storage mcap \
  --storage-preset-profile zstd_fast \
  --max-bag-size 1000000000 \
  --max-bag-duration 300 \
  --qos-profile-overrides-path qos_overrides.yaml \
  -o /data/bags/amr-07_$(date +%Y%m%dT%H%M%S) \
  -e '^/(tf|tf_static|odom|cmd_vel|diagnostics|battery_state|joint_states|nav2/.*)$'

The -e (--regex) flag selects topics by regular expression, which is easier to maintain than long topic lists. Recent distributions also offer exclusion filters (--exclude-regex and --exclude-topics on Jazzy; -x on older releases) for the opposite pattern of "everything except raw images".

Close-up of a single-board computer used as an onboard logging computer on a mobile robot
Onboard computers have limited storage, CPU and I/O bandwidth, so the recorder must run within a strict budget alongside perception and control.

Topic selection policies: always-on vs event-triggered recording

The most effective cost lever in a robot data pipeline is deciding what not to record. Define two or three recording profiles and version them like your navigation parameters.

The always-on profile

Record low-bandwidth, high-value topics continuously: TF, odometry, velocity commands, localization status, diagnostics, battery and joint states, mission state and error events. For a typical mobile robot this is, for example, 50–300 MB per hour after compression: cheap enough to keep for every shift, and enough to answer most "what was the robot doing?" questions.

Event-triggered recording with snapshot mode

Cameras, depth images and point clouds are recorded only around interesting events. rosbag2 snapshot mode keeps messages in an in-memory circular buffer bounded by --max-cache-size and writes it to disk only when the snapshot service is called.

# Keep roughly the last 1.5 GB of camera and lidar data in memory
ros2 bag record --snapshot-mode \
  --storage mcap \
  --max-cache-size 1500000000 \
  -o /data/snapshots/amr-07 \
  /front_camera/image_raw/compressed /lidar/points /tf /odom

# From an event detector, safety node or operator UI:
ros2 service call /rosbag2_recorder/snapshot rosbag2_interfaces/srv/Snapshot

Typical triggers: emergency stops, localization confidence dropping, an aborted navigation goal, a bumper event, an operator pressing "flag this", or an on-robot anomaly score. Because the buffer holds data from before the trigger, you capture the lead-up to the incident. Size it from the data rate: at, for example, 25 MB/s of camera and lidar data, a 1.5 GB cache holds about 60 seconds.

A third, optional profile is campaign recording: full-rate recording for a defined period, for example to collect training data at a new site. Time-box and approve campaigns; they produce orders of magnitude more data. If you run perception on the robot, the same triggers can come from models; see our guide to edge AI video analytics for robots.

Every rosbag2 bag directory contains a metadata.yaml with the storage identifier, start time, duration, files, compression settings, and every topic with its type, QoS and message count:

ros2 bag info /data/bags/amr-07_20261005T140211

That metadata does not know which robot, site, software build or mission produced the bag. Recent rosbag2 versions support a custom_data section (set with --custom-data key=value), or you can write a sidecar JSON file. Either way, attach at least:

  • robot ID, hardware revision and site
  • software version or git SHA of the robot stack, plus the recording profile version
  • mission or job ID, and operator or customer context where permitted
  • trigger type and event ID for snapshot bags
  • a content hash of each file, so later stages can deduplicate and verify integrity

At ingest, parse the MCAP summary section (cheap, since it sits at the end of the file) and write one row per bag and per topic into a searchable index, plus an event table linking incidents to time ranges. Engineers should be able to ask "all front-camera data within 30 seconds of an emergency stop, on firmware 4.2, at site Hamburg, last week" and receive a list of byte ranges rather than a list of 2 GB files. Content hashes also make the log verifiable later; we explain how to turn them into tamper-evident logs in our article on Merkle trees for verifiable machine data.

Upload strategies and retention tiers

Bandwidth-aware, resumable, prioritized uploads

Uploading is where most homegrown pipelines fail. A production uploader should implement:

  • Completed-file detection. Only upload files the recorder has closed.
  • Resumable transfers. Multipart uploads with per-part checksums, so an interrupted 1 GB transfer resumes at the last confirmed part.
  • Priority queues. Incident snapshots first, always-on telemetry second, campaign data last; newest first within a class.
  • Bandwidth shaping. Cap throughput while the robot works; open the throttle when docked or on a known-good network.
  • Disk pressure policy. For example, at 80% stop campaign recording, at 95% delete already-uploaded files, and never delete un-uploaded incident bags.
  • Verification before deletion. Delete locally only after the server confirms the stored hash matches.

Data retention tiers

Not all bags deserve the same storage class. A typical tiering policy looks like this:

TierContentsTypical retentionWhere it lives
HotRecent incidents, last 7–30 days of telemetry, extracted ParquetDays to weeksObject storage close to analytics compute
WarmBags referenced by open tickets, curated training setsMonthsLower-cost object storage or a regional data center
Cold / archiveRaw bags kept for audits, warranty claims and future retrainingYearsArchive tier or decentralized storage network
DeleteUneventful high-bandwidth data with no downstream useNever stored long-termDropped after extraction or summary

Visual data dominates storage cost, so the archive tier is where most savings come from; we break down the numbers in hybrid cloud storage for robotics.

Extracting data: from MCAP bags to images, video and Parquet

Raw bags are the system of record but a poor analytics format. Run extractors once at ingest and store derived artifacts alongside the bag: Parquet tables for numeric topics, JPEG frames or MP4 clips for cameras. The example uses the mcap and mcap-ros2-support Python packages; the rosbags library is a pure-Python alternative that reads sqlite3 and MCAP bags.

from pathlib import Path

import pyarrow as pa
import pyarrow.parquet as pq
from mcap.reader import make_reader
from mcap_ros2.reader import read_ros2_messages

BAG = Path("amr-07_20261005T140211/amr-07_20261005T140211_0.mcap")

# 1. Inspect the summary section: cheap, no full scan required.
with open(BAG, "rb") as f:
    summary = make_reader(f).get_summary()
    for cid, ch in summary.channels.items():
        count = summary.statistics.channel_message_counts.get(cid, 0)
        print(f"{ch.topic:40s} {count:8d} msgs")

# 2. Extract odometry into a columnar table.
cols = {"t_ns": [], "x": [], "y": [], "yaw_z": [], "vx": [], "wz": []}
for m in read_ros2_messages(str(BAG), topics=["/odom"]):
    odom = m.ros_msg
    cols["t_ns"].append(m.log_time_ns)
    cols["x"].append(odom.pose.pose.position.x)
    cols["y"].append(odom.pose.pose.position.y)
    cols["yaw_z"].append(odom.pose.pose.orientation.z)
    cols["vx"].append(odom.twist.twist.linear.x)
    cols["wz"].append(odom.twist.twist.angular.z)

pq.write_table(pa.table(cols), "odom.parquet", compression="zstd")

# 3. Dump compressed camera frames as JPEG files.
out = Path("frames")
out.mkdir(exist_ok=True)
for m in read_ros2_messages(str(BAG), topics=["/front_camera/image_raw/compressed"]):
    (out / f"{m.log_time_ns}.jpg").write_bytes(bytes(m.ros_msg.data))

Frames can then be encoded into MP4 clips with ffmpeg, and Parquet queried with DuckDB, Spark or your warehouse. Partition by date and robot so fleet-wide queries read only what they need. For format migrations, ros2 bag convert rewrites bags with new storage, compression, splitting or topic filters:

# convert.yaml: sqlite3 legacy bag to MCAP, split, telemetry only
output_bags:
  - uri: amr-07_legacy_mcap
    storage_id: mcap
    storage_preset_profile: zstd_small
    max_bagfile_duration: 300
    all_topics: true   # use "all: true" on Humble
ros2 bag convert -i amr-07_legacy_sqlite -o convert.yaml

Reference architecture for a fleet robot data pipeline

Putting the pieces together, a production pipeline has six stages:

 [Robot]                       [Edge / site]            [Cloud / hybrid storage]
 ROS 2 nodes                    Edge buffer               Object storage (hot)
   |  rosbag2 (MCAP,             (NAS / gateway,             |
   |  split, profiles,           store-and-forward)          v
   |  snapshots)                    |                    Ingest + index
   v                                |                    (metadata, events,
 Agent on robot  ---- LAN --------->+---- WAN, resumable,  hashes)
 (watch, hash, queue,                   prioritized ------>  |
  shape bandwidth)                                          v
                                                       Extractors
                                                       (video, Parquet,
                                                        enrichment) --> Warm / cold tiers
  1. Robot. rosbag2 with MCAP, splitting, versioned profiles and a snapshot recorder.
  2. Agent. Watches the bag directory, hashes completed files, enforces disk policy and uploads by priority.
  3. Edge buffer. At sites with many robots or weak WAN, a gateway absorbs LAN bursts and forwards at a controlled rate; a natural place for edge compute such as transcoding or anonymization.
  4. Object storage. Hash-verified storage for raw bags, with lifecycle rules into warm and cold tiers.
  5. Index. A catalog of bags, topics, time ranges, versions and events for search and access control.
  6. Extractors. Containerized jobs producing video, images, Parquet and enriched logs.

This is the architecture behind MerkleBot's data pipelines: the MerkleBot Agent runs on the robot or edge device with out-of-the-box support for ROS, ROS 2 and hardware such as Universal Robots arms and Boston Dynamics Spot, and extraction runs as pre-built ROSbag extractors or your own Docker containers. The REST API and CLI are documented on the developer page.

Production checklist

AreaCheck
FormatMCAP storage everywhere; legacy sqlite3 bags converted at ingest
CompressionMCAP chunk compression via a preset profile; cameras recorded as compressed transport
Splitting--max-bag-size or --max-bag-duration set; files roughly 0.5–2 GB
QoSOverride file for latched and best-effort topics; verified that /tf_static is captured
Topic policyVersioned always-on, snapshot and campaign profiles; regex-based selection
SnapshotsCache size derived from data rate and RAM budget; triggers wired to safety and anomaly events
MetadataRobot, site, software SHA, profile version, event IDs and file hashes attached to every bag
UploadResumable multipart, priority classes, bandwidth caps, hash verification before local delete
DiskExplicit 80% / 95% disk-pressure policy; recorder never starves other processes
RetentionHot, warm, cold and delete rules documented and automated
ExtractionVersioned extractors producing Parquet, frames and clips, partitioned by date and robot
ObservabilityDashboards for bytes recorded, queued and uploaded per robot, plus dropped-message counts

Frequently asked questions

Should I use rosbag2 compression or MCAP compression?

For most fleets, use MCAP's built-in chunk compression through a storage preset profile such as zstd_fast. It compresses well and keeps files seekable and indexable. rosbag2 file-mode compression produces a better ratio only marginally and forces full decompression before reading, so reserve it for data going straight to cold archive.

How large should individual bag files be?

A typical sweet spot is 500 MB to 2 GB, or 1 to 5 minutes per file. Smaller files increase per-file overhead in upload and indexing; larger files make retries, partial downloads and corruption more expensive.

Can I read MCAP files without installing ROS?

Yes. MCAP embeds message schemas, and official libraries exist for Python, C++, Go, Rust and TypeScript. The mcap-ros2-support Python package and the rosbags library both decode ROS 2 messages without a ROS installation, which keeps cloud extractors lightweight.

How do I make sure I capture data from before an incident?

Run a recorder in snapshot mode for high-bandwidth topics. It keeps a circular buffer in memory bounded by --max-cache-size and writes it to disk when the snapshot service is called, so the bag contains the seconds leading up to the trigger. Pair it with an always-on recorder for low-bandwidth telemetry.

Conclusion

ROS 2 data logging at fleet scale is less about the record button and more about policy: which topics, in which format and file size, uploaded in which order, kept for how long and extracted into what. rosbag2 with MCAP is a solid foundation on the robot; the pipeline behind it is what turns thousands of bag files into a dataset your engineers and models can use.

If you would rather not build and operate that pipeline yourself, book a 30-minute demo and we will walk through how MerkleBot connects your ROS 2 fleet, from the first robot to the whole fleet.

MerkleBot EngineeringThe team behind MerkleBot's data platform for robotics and IoT — robotics engineers working on pipelines, hybrid storage and data-driven business models for machines.

Keep reading

Related articles

All articles

Ready to put your machine data to work?

In a 30-minute demo we'll map your fleet's data flows and show where hybrid storage and compute cut costs.

Book a demo