Recording a rosbag on a single robot on your desk is easy. Recording the right data, on fifty robots in three countries, every day, without filling disks, saturating uplinks or losing the one bag that explains last Tuesday's collision, is a data engineering problem. This guide covers how to configure rosbag2 and MCAP for production, and how to build the pipeline that sits behind them.
Key takeaways
- Use the MCAP storage plugin (the rosbag2 default since ROS 2 Iron) and prefer its built-in chunk compression over rosbag2 file-mode compression, so bags stay seekable.
- Split bags by size or duration so that every file is a small, independently uploadable and independently useful unit.
- Combine a lean always-on recording profile with event-triggered snapshots from an in-memory circular buffer for high-bandwidth sensors.
- Treat upload as a scheduling problem: bandwidth-aware, resumable, and prioritized by incident severity rather than by file age.
- Index bags by robot, time, topic and event at ingest time, and extract to columnar formats like Parquet so analytics never has to open a raw bag.
Why ROS 2 data logging breaks at fleet scale
Most teams start ROS 2 data logging the same way: an engineer runs ros2 bag record -a, copies the result off the robot with scp, and opens it in a visualizer. That workflow is fine for a prototype. It fails at fleet scale for predictable reasons:
- Volume explodes. A mobile robot with four compressed cameras, a 3D lidar and the usual odometry, TF and diagnostics can, for example, produce tens of gigabytes per operating hour. Across a fleet, recording everything stops being an option.
- Disks fill silently. A recorder that runs until a 256 GB embedded disk is full can take down other processes, including the ones writing the logs you need for the post-mortem.
- Uplinks are unreliable. Warehouse Wi-Fi, field LTE and strict customer networks interrupt uploads. Without resumable transfers you resend the same gigabytes repeatedly.
- Nobody can find anything. Ten thousand bags named by timestamp in a bucket are not a dataset.
- Recording affects the robot. Wrong QoS or CPU and I/O contention with perception nodes can drop messages or add control-loop latency.
A robust robot data pipeline therefore has two halves: rosbag2 configured to record the right topics, in the right format, in bounded space; and a fleet data management layer that moves, indexes, tiers and extracts that data.
rosbag2 storage plugins: sqlite3 vs MCAP
rosbag2 separates the recorder from the on-disk format through storage plugins. Two plugins matter in practice.
sqlite3 was the original default: each message is a row in a SQLite database. It is robust, but it carries per-message overhead, is not designed for streaming large binary payloads, and does not store the message definitions needed to decode it outside ROS.
MCAP is an open, self-describing container for timestamped, multi-channel messages, and the default rosbag2 storage since ROS 2 Iron. Its properties fit robotics well:
- Self-describing. Message schemas are embedded, so a bag decodes years later without the original workspace.
- Chunked and indexed. A summary section at the end of the file enables fast seeking and per-topic reads without a full scan.
- Append-friendly. Built for streaming writes; a file truncated by power loss is usually recoverable up to the last complete chunk.
- Language-neutral. Libraries exist for Python, C++, Go, Rust and TypeScript, so cloud extractors need no ROS installation.
The specification is documented at mcap.dev. On Humble, install the rosbag2_storage_mcap package; on any distribution, pass --storage mcap explicitly in launch files to avoid surprises.
Configuring rosbag2 for production: compression, splitting and QoS
The defaults of ros2 bag record are designed for interactive use. For production, three settings matter most.
Compression: zstd, file mode vs message mode vs MCAP chunks
rosbag2 has its own zstd compression layer with two modes. File mode compresses each completed file as a whole: good ratio, but the result must be decompressed before you can seek inside it. Message mode compresses each message individually: seekable, but small messages compress poorly.
With MCAP there is a better option: the storage plugin compresses whole chunks while keeping the index readable. Select it with a storage preset profile such as zstd_fast (lower CPU) or zstd_small (better ratio), or fastwrite to skip chunk compression when CPU is the bottleneck. Reserve rosbag2 file mode for bags going straight to cold storage.
Be realistic: structured topics such as odometry, TF and joint states typically compress 3–10x, while JPEG or H.264 payloads barely shrink. If you record raw sensor_msgs/Image, the biggest win is recording a compressed transport instead.
Splitting bags by size and duration
Never record one giant file per session. Use --max-bag-size (bytes) or --max-bag-duration (seconds) so the recorder rolls over regularly. Small files are easier to upload, retry and delete, and a corrupted file costs minutes of data rather than hours. Files of roughly 500 MB to 2 GB, or 1–5 minutes, are a good balance.
QoS overrides
The recorder adapts to publisher QoS, but be explicit for latched topics such as /tf_static (which need transient_local durability to capture messages published before recording started) and for high-rate sensors where best-effort avoids back-pressure:
# qos_overrides.yaml
/tf_static:
history: keep_all
reliability: reliable
durability: transient_local
/robot_description:
history: keep_last
depth: 1
reliability: reliable
durability: transient_local
/lidar/points:
history: keep_last
depth: 10
reliability: best_effort
durability: volatile
Putting it together, a production recording command looks like this:
ros2 bag record \
--storage mcap \
--storage-preset-profile zstd_fast \
--max-bag-size 1000000000 \
--max-bag-duration 300 \
--qos-profile-overrides-path qos_overrides.yaml \
-o /data/bags/amr-07_$(date +%Y%m%dT%H%M%S) \
-e '^/(tf|tf_static|odom|cmd_vel|diagnostics|battery_state|joint_states|nav2/.*)$'
The -e (--regex) flag selects topics by regular expression, which is easier to maintain than long topic lists. Recent distributions also offer exclusion filters (--exclude-regex and --exclude-topics on Jazzy; -x on older releases) for the opposite pattern of "everything except raw images".

Topic selection policies: always-on vs event-triggered recording
The most effective cost lever in a robot data pipeline is deciding what not to record. Define two or three recording profiles and version them like your navigation parameters.
The always-on profile
Record low-bandwidth, high-value topics continuously: TF, odometry, velocity commands, localization status, diagnostics, battery and joint states, mission state and error events. For a typical mobile robot this is, for example, 50–300 MB per hour after compression: cheap enough to keep for every shift, and enough to answer most "what was the robot doing?" questions.
Event-triggered recording with snapshot mode
Cameras, depth images and point clouds are recorded only around interesting events. rosbag2 snapshot mode keeps messages in an in-memory circular buffer bounded by --max-cache-size and writes it to disk only when the snapshot service is called.
# Keep roughly the last 1.5 GB of camera and lidar data in memory
ros2 bag record --snapshot-mode \
--storage mcap \
--max-cache-size 1500000000 \
-o /data/snapshots/amr-07 \
/front_camera/image_raw/compressed /lidar/points /tf /odom
# From an event detector, safety node or operator UI:
ros2 service call /rosbag2_recorder/snapshot rosbag2_interfaces/srv/Snapshot
Typical triggers: emergency stops, localization confidence dropping, an aborted navigation goal, a bumper event, an operator pressing "flag this", or an on-robot anomaly score. Because the buffer holds data from before the trigger, you capture the lead-up to the incident. Size it from the data rate: at, for example, 25 MB/s of camera and lidar data, a 1.5 GB cache holds about 60 seconds.
A third, optional profile is campaign recording: full-rate recording for a defined period, for example to collect training data at a new site. Time-box and approve campaigns; they produce orders of magnitude more data. If you run perception on the robot, the same triggers can come from models; see our guide to edge AI video analytics for robots.
Metadata, indexing and search
Every rosbag2 bag directory contains a metadata.yaml with the storage identifier, start time, duration, files, compression settings, and every topic with its type, QoS and message count:
ros2 bag info /data/bags/amr-07_20261005T140211
That metadata does not know which robot, site, software build or mission produced the bag. Recent rosbag2 versions support a custom_data section (set with --custom-data key=value), or you can write a sidecar JSON file. Either way, attach at least:
- robot ID, hardware revision and site
- software version or git SHA of the robot stack, plus the recording profile version
- mission or job ID, and operator or customer context where permitted
- trigger type and event ID for snapshot bags
- a content hash of each file, so later stages can deduplicate and verify integrity
At ingest, parse the MCAP summary section (cheap, since it sits at the end of the file) and write one row per bag and per topic into a searchable index, plus an event table linking incidents to time ranges. Engineers should be able to ask "all front-camera data within 30 seconds of an emergency stop, on firmware 4.2, at site Hamburg, last week" and receive a list of byte ranges rather than a list of 2 GB files. Content hashes also make the log verifiable later; we explain how to turn them into tamper-evident logs in our article on Merkle trees for verifiable machine data.
Upload strategies and retention tiers
Bandwidth-aware, resumable, prioritized uploads
Uploading is where most homegrown pipelines fail. A production uploader should implement:
- Completed-file detection. Only upload files the recorder has closed.
- Resumable transfers. Multipart uploads with per-part checksums, so an interrupted 1 GB transfer resumes at the last confirmed part.
- Priority queues. Incident snapshots first, always-on telemetry second, campaign data last; newest first within a class.
- Bandwidth shaping. Cap throughput while the robot works; open the throttle when docked or on a known-good network.
- Disk pressure policy. For example, at 80% stop campaign recording, at 95% delete already-uploaded files, and never delete un-uploaded incident bags.
- Verification before deletion. Delete locally only after the server confirms the stored hash matches.
Data retention tiers
Not all bags deserve the same storage class. A typical tiering policy looks like this:
| Tier | Contents | Typical retention | Where it lives |
|---|---|---|---|
| Hot | Recent incidents, last 7–30 days of telemetry, extracted Parquet | Days to weeks | Object storage close to analytics compute |
| Warm | Bags referenced by open tickets, curated training sets | Months | Lower-cost object storage or a regional data center |
| Cold / archive | Raw bags kept for audits, warranty claims and future retraining | Years | Archive tier or decentralized storage network |
| Delete | Uneventful high-bandwidth data with no downstream use | Never stored long-term | Dropped after extraction or summary |
Visual data dominates storage cost, so the archive tier is where most savings come from; we break down the numbers in hybrid cloud storage for robotics.
Extracting data: from MCAP bags to images, video and Parquet
Raw bags are the system of record but a poor analytics format. Run extractors once at ingest and store derived artifacts alongside the bag: Parquet tables for numeric topics, JPEG frames or MP4 clips for cameras. The example uses the mcap and mcap-ros2-support Python packages; the rosbags library is a pure-Python alternative that reads sqlite3 and MCAP bags.
from pathlib import Path
import pyarrow as pa
import pyarrow.parquet as pq
from mcap.reader import make_reader
from mcap_ros2.reader import read_ros2_messages
BAG = Path("amr-07_20261005T140211/amr-07_20261005T140211_0.mcap")
# 1. Inspect the summary section: cheap, no full scan required.
with open(BAG, "rb") as f:
summary = make_reader(f).get_summary()
for cid, ch in summary.channels.items():
count = summary.statistics.channel_message_counts.get(cid, 0)
print(f"{ch.topic:40s} {count:8d} msgs")
# 2. Extract odometry into a columnar table.
cols = {"t_ns": [], "x": [], "y": [], "yaw_z": [], "vx": [], "wz": []}
for m in read_ros2_messages(str(BAG), topics=["/odom"]):
odom = m.ros_msg
cols["t_ns"].append(m.log_time_ns)
cols["x"].append(odom.pose.pose.position.x)
cols["y"].append(odom.pose.pose.position.y)
cols["yaw_z"].append(odom.pose.pose.orientation.z)
cols["vx"].append(odom.twist.twist.linear.x)
cols["wz"].append(odom.twist.twist.angular.z)
pq.write_table(pa.table(cols), "odom.parquet", compression="zstd")
# 3. Dump compressed camera frames as JPEG files.
out = Path("frames")
out.mkdir(exist_ok=True)
for m in read_ros2_messages(str(BAG), topics=["/front_camera/image_raw/compressed"]):
(out / f"{m.log_time_ns}.jpg").write_bytes(bytes(m.ros_msg.data))
Frames can then be encoded into MP4 clips with ffmpeg, and Parquet queried with DuckDB, Spark or your warehouse. Partition by date and robot so fleet-wide queries read only what they need. For format migrations, ros2 bag convert rewrites bags with new storage, compression, splitting or topic filters:
# convert.yaml: sqlite3 legacy bag to MCAP, split, telemetry only
output_bags:
- uri: amr-07_legacy_mcap
storage_id: mcap
storage_preset_profile: zstd_small
max_bagfile_duration: 300
all_topics: true # use "all: true" on Humble
ros2 bag convert -i amr-07_legacy_sqlite -o convert.yaml
Reference architecture for a fleet robot data pipeline
Putting the pieces together, a production pipeline has six stages:
[Robot] [Edge / site] [Cloud / hybrid storage]
ROS 2 nodes Edge buffer Object storage (hot)
| rosbag2 (MCAP, (NAS / gateway, |
| split, profiles, store-and-forward) v
| snapshots) | Ingest + index
v | (metadata, events,
Agent on robot ---- LAN --------->+---- WAN, resumable, hashes)
(watch, hash, queue, prioritized ------> |
shape bandwidth) v
Extractors
(video, Parquet,
enrichment) --> Warm / cold tiers
- Robot. rosbag2 with MCAP, splitting, versioned profiles and a snapshot recorder.
- Agent. Watches the bag directory, hashes completed files, enforces disk policy and uploads by priority.
- Edge buffer. At sites with many robots or weak WAN, a gateway absorbs LAN bursts and forwards at a controlled rate; a natural place for edge compute such as transcoding or anonymization.
- Object storage. Hash-verified storage for raw bags, with lifecycle rules into warm and cold tiers.
- Index. A catalog of bags, topics, time ranges, versions and events for search and access control.
- Extractors. Containerized jobs producing video, images, Parquet and enriched logs.
This is the architecture behind MerkleBot's data pipelines: the MerkleBot Agent runs on the robot or edge device with out-of-the-box support for ROS, ROS 2 and hardware such as Universal Robots arms and Boston Dynamics Spot, and extraction runs as pre-built ROSbag extractors or your own Docker containers. The REST API and CLI are documented on the developer page.
Production checklist
| Area | Check |
|---|---|
| Format | MCAP storage everywhere; legacy sqlite3 bags converted at ingest |
| Compression | MCAP chunk compression via a preset profile; cameras recorded as compressed transport |
| Splitting | --max-bag-size or --max-bag-duration set; files roughly 0.5–2 GB |
| QoS | Override file for latched and best-effort topics; verified that /tf_static is captured |
| Topic policy | Versioned always-on, snapshot and campaign profiles; regex-based selection |
| Snapshots | Cache size derived from data rate and RAM budget; triggers wired to safety and anomaly events |
| Metadata | Robot, site, software SHA, profile version, event IDs and file hashes attached to every bag |
| Upload | Resumable multipart, priority classes, bandwidth caps, hash verification before local delete |
| Disk | Explicit 80% / 95% disk-pressure policy; recorder never starves other processes |
| Retention | Hot, warm, cold and delete rules documented and automated |
| Extraction | Versioned extractors producing Parquet, frames and clips, partitioned by date and robot |
| Observability | Dashboards for bytes recorded, queued and uploaded per robot, plus dropped-message counts |
Frequently asked questions
Should I use rosbag2 compression or MCAP compression?
For most fleets, use MCAP's built-in chunk compression through a storage preset profile such as zstd_fast. It compresses well and keeps files seekable and indexable. rosbag2 file-mode compression produces a better ratio only marginally and forces full decompression before reading, so reserve it for data going straight to cold archive.
How large should individual bag files be?
A typical sweet spot is 500 MB to 2 GB, or 1 to 5 minutes per file. Smaller files increase per-file overhead in upload and indexing; larger files make retries, partial downloads and corruption more expensive.
Can I read MCAP files without installing ROS?
Yes. MCAP embeds message schemas, and official libraries exist for Python, C++, Go, Rust and TypeScript. The mcap-ros2-support Python package and the rosbags library both decode ROS 2 messages without a ROS installation, which keeps cloud extractors lightweight.
How do I make sure I capture data from before an incident?
Run a recorder in snapshot mode for high-bandwidth topics. It keeps a circular buffer in memory bounded by --max-cache-size and writes it to disk when the snapshot service is called, so the bag contains the seconds leading up to the trigger. Pair it with an always-on recorder for low-bandwidth telemetry.
Conclusion
ROS 2 data logging at fleet scale is less about the record button and more about policy: which topics, in which format and file size, uploaded in which order, kept for how long and extracted into what. rosbag2 with MCAP is a solid foundation on the robot; the pipeline behind it is what turns thousands of bag files into a dataset your engineers and models can use.
If you would rather not build and operate that pipeline yourself, book a 30-minute demo and we will walk through how MerkleBot connects your ROS 2 fleet, from the first robot to the whole fleet.





