Storage & Compute

Hybrid Cloud Storage for Robotics: How to Cut Visual Data Costs Up to 10x

Edge compute module with a large heatsink used to pre-process robot video

For most robotics companies, the cloud bill grows faster than the fleet. The culprit is rarely compute; it is the steady accumulation of camera, depth and lidar data that nobody wants to delete and few people actually read. This article shows where robotics data storage costs come from, and how a hybrid cloud storage design with edge reduction and a decentralized archive can cut the cost of visual data by up to 10x.

Key takeaways

  • Visual sensors dominate data volume: one raw 1080p camera produces more data in an hour than a robot's structured telemetry produces in months.
  • Storage price per TB is only one cost driver; egress, API requests, retrieval fees and minimum storage durations often matter as much.
  • The cheapest byte is the one you never upload: event-based capture, downsampling and deduplication at the edge reduce every downstream cost.
  • Content-addressed, decentralized storage is well suited to large, immutable video archives, provided you encrypt before upload and plan for retrieval latency.
  • A hybrid design keeps hot analytics on a hyperscaler, warm data in a local data center and the long-tail archive on a decentralized network.

Why robots generate so much data

A robot is a bundle of sensors attached to a computer, and most of those sensors produce images. Structured telemetry such as odometry, joint states or battery levels is tiny by comparison. The table below gives illustrative data rates for common sensors; real numbers depend on resolution, frame rate, encoding and message overhead.

Sensor (illustrative configuration)CalculationApprox. ratePer hour
RGB camera, 1920×1080, 30 fps, raw 8-bit RGB1920 × 1080 × 3 B × 30~187 MB/s~670 GB
Same camera, H.264 at about 8 Mbit/s8 Mbit/s ÷ 8~1 MB/s~3.6 GB
Depth camera, 640×480, 16-bit, 30 fps640 × 480 × 2 B × 30~18 MB/s~66 GB
3D lidar, 64 channels × 1024 points, 10 Hz, ~16 B/point655,360 pts/s × 16 B~10.5 MB/s~38 GB
IMU at 400 Hz, ~300 B per ROS message400 × 300 B~0.12 MB/s~0.4 GB
Joint states, 6 axes at 500 Hz, ~200 B per message500 × 200 B~0.1 MB/s~0.36 GB

Even with every camera encoded as H.264, a mobile robot with four cameras, a depth sensor recorded at reduced rate and a lidar can produce, for example, 40–80 GB per operating hour. A fleet of 50 robots working two shifts would generate tens of terabytes per day if it uploaded everything. No team does that for long, but even a disciplined subset adds up to hundreds of terabytes per year, and visual data is typically 90% or more of it.

Unlike logs, this data is valuable for years: it is the raw material for retraining perception models, reproducing field incidents, defending warranty and liability claims and building simulation scenarios. So the default becomes "keep everything", and the storage bill compounds month after month.

Where robotics data storage costs come from

Cloud cost optimization starts with an honest bill breakdown. For robotics workloads, five line items usually matter:

  • Capacity. The per-TB-month price of stored data. It scales with retention, so a 12-month retention policy means you pay for 12 months of ingest at steady state.
  • Egress. Data transferred out of the provider's network, for example to an on-premises GPU cluster, a labelling vendor or another cloud. For training-heavy teams, egress can rival capacity.
  • Requests. Per-operation charges for PUT, GET and LIST calls. Millions of small files, such as individual JPEG frames, can make request costs surprisingly large.
  • Retrieval and restore. Archive tiers are cheap to hold but charge per GB to read back, may take hours to restore and often have minimum storage durations with early-deletion fees.
  • Ingest infrastructure. Gateways, VPNs, transfer services and the engineering time to operate them.

Hot, warm and cold tiers

Tiering matches the price of storage to the probability that data will be read:

TierTypical robotics contentAccess patternGood fit
HotLast 7–30 days, open incidents, active training sets, extracted ParquetFrequent, low latency, often from cloud computeHyperscaler object storage next to analytics
WarmLast 1–6 months, curated datasets, reference runsOccasional, minutes are acceptableRegional or local data center, on-premises NAS
Cold / archiveRaw video and point clouds kept for yearsRare, bulk, latency of hours acceptableArchive tiers or decentralized storage networks

The key observation is that the bulk of a robotics video data archive sits in the cold tier, and that is where pricing differences between providers and architectures are largest.

Embedded board with an SD card slot used for local data storage on a robot or edge device
Local storage on the robot or edge gateway is the first tier: it buffers data until it can be reduced, prioritized and uploaded.

Edge pre-processing and data reduction

Every byte removed at the edge saves capacity, transfer, requests and retrieval in every tier after it. Data reduction should be the first step of any cloud cost optimization project, before negotiating storage prices.

  • Encode, do not store raw. Record cameras as H.264, H.265 or AV1 instead of raw frames. For perception training, a high-quality encoding preset is usually sufficient and orders of magnitude smaller.
  • Event-based capture. Keep high-bandwidth sensors in a rolling buffer and persist only windows around events such as e-stops, anomalies, interventions or operator flags. A 60-second window per event instead of continuous recording can, for example, reduce camera volume by 80–95%.
  • Downsample by purpose. Navigation debugging rarely needs 30 fps; 5 fps thumbnails often suffice for search and review, with full rate only around events.
  • Drop redundant data. A robot parked at a charger for six hours produces six hours of nearly identical frames. Motion-gated recording or perceptual-hash deduplication removes them.
  • Batch small objects. Package frames and small files into larger containers, such as MCAP files or CAR archives, to cut request costs and metadata overhead.
  • Compress structured data. zstd on telemetry typically yields 3–10x, cheaply.

We cover the recording side in detail in our guide to ROS 2 data logging with rosbag2 and MCAP. On the infrastructure side, the reduction steps run best on an edge compute node at the site, so robots stay focused on their primary task.

Content-addressed and decentralized storage

Traditional object storage is location-addressed: you ask a specific provider for a specific bucket and key. Content-addressed storage identifies data by a cryptographic hash of its contents instead. That small change has large consequences for robotics archives.

CIDs and IPFS

In IPFS, data is split into blocks, arranged into a Merkle DAG, and identified by a content identifier (CID), which encodes the hash of the root of that structure plus format information. The IPFS documentation on content addressing describes the details. For an archive this means:

  • Integrity by construction. If the bytes you get back hash to the CID you stored, they are the bytes you uploaded, regardless of who served them.
  • Automatic deduplication. Identical content yields identical CIDs, so re-uploads of the same file cost nothing extra.
  • Provider independence. The same CID can be served by any node that holds the data, which avoids lock-in to a single provider's namespace.

IPFS on its own does not guarantee persistence: data stays available only while some node pins it. Long-term persistence needs a storage layer with explicit commitments.

Filecoin storage deals

Filecoin adds that layer. A client makes storage deals with independent storage providers to keep data for an agreed duration. Providers must regularly submit cryptographic proofs, Proof-of-Replication when data is sealed and Proof-of-Spacetime over the deal's lifetime, which are verified on a public chain. See the Filecoin documentation for the protocol specifics. The practical properties are:

  • Redundancy. Store each archive with several independent providers in different regions, for example three to five replicas, so no single company, data center or jurisdiction is a single point of failure.
  • Verifiability. Storage is continuously proven rather than promised in a contract.
  • Retrievability. Retrieval speed varies by provider. Sealed data may need to be unsealed before it can be served, so for data you might need within minutes, choose providers that keep fast-retrieval copies or keep a cache in a warmer tier.

Decentralized storage is not a fit for every workload. It works best for exactly what robotics fleets have in abundance: large, immutable, rarely read files such as video, point clouds and completed rosbags.

A hybrid cloud storage architecture for robotics

Hybrid cloud storage combines each tier where it is strongest:

  1. Robot and edge. Recording with bounded local buffers, event-based capture, encoding, deduplication and encryption. An agent on the robot hashes and queues completed files.
  2. Local or regional data center. A warm tier for recent months of data, close to on-premises GPU training clusters so reads incur no hyperscaler egress.
  3. Decentralized archive. Encrypted, content-addressed copies of everything worth keeping long-term, replicated across independent providers.
  4. Hyperscaler hot tier. Only the data that cloud analytics, dashboards and managed ML services need right now, with short lifecycle rules.
  5. Catalog. One index mapping logical datasets (robot, time, topic, event) to physical locations and CIDs, so users search by meaning, not by tier.

The catalog is what makes the architecture usable. Engineers should never need to know whether a clip currently lives in the warm tier or the archive; the platform resolves it and restores it if necessary. This is how MerkleBot hybrid storage is built: edge compute, local data center providers and decentralized storage networks behind a single API, which removes single points of failure. One mobile robotics customer saved up to 10x on storing large video and visual datasets by backing up the archive to a decentralized storage network.

A worked cost model: all-hyperscaler vs hybrid

The following model is an illustrative estimate. All prices are assumed round numbers chosen for readability, not quotes from any provider; plug in your own contract rates before drawing conclusions.

Assumptions

ParameterAssumed value
Fleet50 mobile robots
Uploaded volume without edge reduction60 TB per month
Retention12 months (steady state)
Edge reduction in hybrid scenario50% (event-based capture, dedup, downsampling)
Data read for training and review per month10% of stored volume
Hyperscaler standard object storageassumed $20 per TB-month
Hyperscaler egressassumed $70 per TB
Local data center storage (warm)assumed $7 per TB-month, bandwidth included
Decentralized archive incl. replicationassumed $2 per TB-month
Reads from warm or archive tiersassumed $10 per TB
Edge gateways and operations (hybrid only)assumed $1,500 per month

Monthly cost at steady state

Line itemA: all in hyperscalerB: hybrid
Hot storage720 TB × $20 = $14,40030 TB (30 days) × $20 = $600
Warm storagen/a90 TB (90 days) × $7 = $630
Archive storageincluded above360 TB (12 months) × $2 = $720
Reads / egress72 TB × $70 = $5,04036 TB × $10 = $360
Requests and API$300$100
Gateways and operationsn/a$1,500
Total per month$19,740$3,910

In this scenario the hybrid design costs about one fifth of the all-hyperscaler baseline. It is worth being precise about where the savings come from:

  • Archive capacity: keeping 360 TB of archive at an assumed $2 instead of $20 per TB-month is a 10x difference on that line alone ($720 vs $7,200).
  • Edge reduction: halving the volume halves every other line, independent of where data is stored.
  • Egress avoidance: training reads served from a local data center next to the GPUs avoid hyperscaler egress entirely.

Two honest caveats. First, if scenario A used lifecycle rules into a hyperscaler archive class, its capacity line would fall substantially, although retrieval fees, restore delays and minimum durations would then apply to every training read. Second, the hybrid design adds operational complexity, represented here by a flat $1,500; underestimating that line is the most common modelling error. Run the model with your own numbers, and see our pricing page for how MerkleBot packages storage and pipelines.

Risks and how to mitigate them

RiskMitigation
Retrieval latency from the archiveKeep thumbnails, metadata and recent data in warm or hot tiers; choose providers with fast retrieval for data likely to be needed; restore datasets in bulk ahead of training runs
Confidentiality on public networksClient-side encryption (for example AES-256-GCM) before upload; never publish plaintext CIDs of sensitive content
Key managementPer-dataset data keys wrapped by a KMS or HSM-held master key; key rotation and escrow procedures; losing keys means losing data, so test recovery
Compliance and data sovereigntySelect storage providers by region; keep personal data in jurisdiction-specific tiers; anonymize faces and license plates at the edge where required
Right to erasureCrypto-shredding: delete the per-dataset key, rendering all replicas unreadable
Provider churn or failureMultiple independent replicas; automated deal renewal and repair; periodic retrieval tests
Catalog becomes a single point of failureReplicate and back up the catalog; store CIDs and dataset manifests in the archive itself

Migration plan and KPIs

Step-by-step migration plan

  1. Measure. Break down the last three months of storage spend by capacity, egress, requests and retrieval, and by data type. Identify the share of visual data.
  2. Classify. Tag datasets by access frequency, sensitivity and retention requirement. Agree on retention rules with legal, safety and ML teams.
  3. Reduce at the edge. Introduce event-based capture and encoding for new data first. This delivers savings before any migration.
  4. Build the catalog. Index existing data with logical keys and content hashes, so physical moves become invisible to users.
  5. Pilot the archive. Encrypt and archive one cold dataset, for example last year's raw video, with multiple replicas. Measure upload throughput, retrieval time and cost.
  6. Test restores. Restore a full training set from the archive and verify hashes end to end before deleting anything from the original tier.
  7. Migrate in waves. Move cold data by age, then switch lifecycle rules so new data flows hot to warm to archive automatically.
  8. Decommission. Remove redundant copies only after the verification and restore tests pass.

KPIs to track

  • Total storage cost per robot per month, and per operating hour
  • Bytes uploaded per robot-hour, before and after edge reduction
  • Share of data by tier, and share of data never read after 90 days
  • Egress and retrieval cost as a percentage of the storage bill
  • Median and 95th-percentile time to restore a dataset from archive
  • Number of verified replicas per archived dataset, and failed retrieval tests

Frequently asked questions

Is decentralized storage reliable enough for production data?

For immutable archives, yes, when designed properly: store several encrypted replicas with independent providers in different regions, monitor deals and proofs, and run regular retrieval tests. Keep data you need within minutes in a warm or hot tier.

Does the 10x saving apply to my whole cloud bill?

Usually not. Up to 10x applies to the archive storage of large visual datasets. The total saving depends on how much of your bill is archive capacity versus compute, egress and hot storage. In the illustrative model above, the total dropped about 5x.

Can I still run cloud analytics and ML on archived data?

Yes. Restore the relevant subset to the hot tier or to a GPU cluster near your warm tier before the job runs. A catalog that maps datasets to locations makes this a single request rather than a manual hunt.

How do I handle GDPR and data sovereignty?

Anonymize at the edge where possible, encrypt before upload, select storage providers by jurisdiction, and use per-dataset keys so deleting a key renders all replicas unreadable. Document these controls for your data protection officer.

Conclusion

Robotics data storage becomes expensive because visual data is huge, valuable for years and stored by default in the most expensive tier. Reducing data at the edge, tiering it by access pattern and moving the long-tail video data archive to encrypted, content-addressed decentralized storage changes the economics, while a hyperscaler remains the right home for hot analytics. For examples of how fleets put this into practice, see our use cases.

Want to see what hybrid storage would do to your bill? Book a 30-minute demo and bring last quarter's numbers; we will model it with you.

MerkleBot EngineeringThe team behind MerkleBot's data platform for robotics and IoT — robotics engineers working on pipelines, hybrid storage and data-driven business models for machines.

Keep reading

Related articles

All articles

Ready to put your machine data to work?

In a 30-minute demo we'll map your fleet's data flows and show where hybrid storage and compute cut costs.

Book a demo