Machine data is increasingly used as evidence: to settle a warranty claim, price an insurance policy, bill a robot-as-a-service contract or prove where a training dataset came from. Evidence is only useful if the other party can verify it has not been changed. Merkle trees are the simple, decades-old data structure that makes that verification cheap, and this article shows how to use them to build tamper-evident logs for robots and IoT devices.
Key takeaways
- A Merkle tree commits to an arbitrarily large set of records with a single 32-byte root hash; changing any record changes the root.
- Inclusion proofs let a verifier check one record against the root with only about log2(n) hashes: 20 hashes for a million records.
- Append-only Merkle logs with consistency proofs, as standardized for Certificate Transparency in RFC 6962 and RFC 9162, prove that history was only extended, never rewritten.
- Integrity becomes authenticity when roots are signed with a device key held in a secure element or TPM, and time-bound when roots are anchored to an external timestamp.
- Content addressing (CIDs in IPFS and Filecoin) applies the same idea to files, so integrity checks travel with the data wherever it is stored.
The trust problem with machine data
Most robot and IoT data pipelines are built for engineers who trust the data because they own the whole system. That assumption breaks as soon as data crosses an organizational boundary:
- Audits. A safety auditor asks for the logs from the hour before an incident. How do you show they were not edited afterwards?
- Insurance and warranty. A manufacturer denies a claim because an arm exceeded its rated payload; the customer disputes the torque logs.
- SLAs. An uptime guarantee is only as credible as the telemetry that measures it.
- Financing. In robot-as-a-service and data-driven leasing, payments depend on operating hours or cycles, so both parties must trust the counter.
- AI training provenance. When a model fails, you need to prove which data it was trained on and that it was not silently modified or poisoned.
The common requirement is not secrecy but data integrity: anyone holding the data can detect whether it differs from what the device recorded. Tamper-evident is not tamper-proof; edits remain possible but become detectable.
Hashing basics: SHA-256 and BLAKE3
A cryptographic hash function maps input of any length to a fixed-size digest such that it is infeasible to find two inputs with the same digest (collision resistance) or an input matching a given digest (preimage and second-preimage resistance).
- SHA-256 produces a 32-byte digest, is universally supported and hardware-accelerated on most modern CPUs and many microcontrollers. It is the conservative default, and the hash used by Certificate Transparency.
- BLAKE3 also produces 32 bytes by default and is considerably faster in software, especially on large inputs, because it is itself a Merkle tree over 1 KiB chunks and can hash in parallel. It is a strong choice for hashing large files such as video or point clouds on the edge.
Record the algorithm in your data format so you can migrate later, and avoid MD5 and SHA-1, which have practical collision attacks.
How a Merkle tree works
A flat list of record hashes would detect tampering, but proving one record would require sending the whole list. A Merkle tree instead hashes pairwise up to a single root:
root = H(0x01 || h01 || h23)
/ \
h01 = H(0x01 || h0 || h1) h23 = H(0x01 || h2 || h3)
/ \ / \
h0 = H(0x00 || d0) h1 = H(0x00 || d1) h2 = H(0x00 || d2) h3 = H(0x00 || d3)
| | | |
d0 d1 d2 d3
(odometry batch) (joint temps) (e-stop event) (video chunk hash)
- Leaves are hashes of individual records: a telemetry message, a batch of messages, or the hash of a larger file such as a video segment.
- Internal nodes are hashes of their two children concatenated.
- The root commits to every leaf and to their order. Change one bit in
d2andh2,h23and the root all change.
The 0x00 and 0x01 prefixes are domain separation, as specified in RFC 6962: they make it impossible to pass off an internal node as a leaf, which closes a classic second-preimage attack on naive Merkle trees. When the number of leaves is not a power of two, RFC 6962 splits the leaves at the largest power of two smaller than n, so the tree is well defined for any size without padding or duplicating leaves.

Inclusion proofs: verifying one record in O(log n)
To prove that d2 is part of the tree above, you do not need the other records. You only need the sibling hashes along the path to the root: h3 and h01. The verifier computes h2 from d2, combines it with h3 to get h23, combines that with h01 to get the root, and compares with the trusted root.
That path is the inclusion proof (RFC 6962 calls it an audit path), and its length is about log2(n):
| Records in tree | Hashes in proof | Proof size (SHA-256) |
|---|---|---|
| 1,000 | 10 | 320 bytes |
| 1,000,000 | 20 | 640 bytes |
| 1,000,000,000 | 30 | 960 bytes |
An insurer can therefore verify one incident record from a year of fleet telemetry with less than a kilobyte of proof.
Python example: building a root and verifying a proof
The following self-contained example implements RFC 6962-style hashing, builds a root over JSON telemetry records, generates an inclusion proof and verifies it with the verification algorithm from RFC 9162. It uses only the standard library.
import hashlib
import json
def leaf_hash(data: bytes) -> bytes:
return hashlib.sha256(b"\x00" + data).digest()
def node_hash(left: bytes, right: bytes) -> bytes:
return hashlib.sha256(b"\x01" + left + right).digest()
def split_point(n: int) -> int:
"""Largest power of two strictly smaller than n (n > 1)."""
k = 1
while k * 2 < n:
k *= 2
return k
def merkle_root(leaves: list[bytes]) -> bytes:
n = len(leaves)
if n == 0:
return hashlib.sha256(b"").digest()
if n == 1:
return leaf_hash(leaves[0])
k = split_point(n)
return node_hash(merkle_root(leaves[:k]), merkle_root(leaves[k:]))
def inclusion_proof(index: int, leaves: list[bytes]) -> list[bytes]:
n = len(leaves)
if n <= 1:
return []
k = split_point(n)
if index < k:
return inclusion_proof(index, leaves[:k]) + [merkle_root(leaves[k:])]
return inclusion_proof(index - k, leaves[k:]) + [merkle_root(leaves[:k])]
def verify_inclusion(leaf: bytes, index: int, size: int,
proof: list[bytes], root: bytes) -> bool:
if index >= size:
return False
fn, sn, r = index, size - 1, leaf_hash(leaf)
for p in proof:
if sn == 0:
return False
if fn & 1 or fn == sn:
r = node_hash(p, r)
while not fn & 1 and fn != 0:
fn >>= 1
sn >>= 1
else:
r = node_hash(r, p)
fn >>= 1
sn >>= 1
return sn == 0 and r == root
records = [
{"t": 1759650000.120, "robot": "amr-07", "topic": "/battery", "soc": 0.81},
{"t": 1759650000.370, "robot": "amr-07", "topic": "/odom", "x": 12.4, "y": 3.1},
{"t": 1759650001.005, "robot": "amr-07", "topic": "/estop", "state": False},
{"t": 1759650001.250, "robot": "amr-07", "topic": "/joint_temp", "j3": 47.5},
{"t": 1759650002.010, "robot": "amr-07", "topic": "/error", "code": "E_SLIP"},
]
# Canonical serialization: the same record must always produce the same bytes.
leaves = [json.dumps(r, sort_keys=True, separators=(",", ":")).encode() for r in records]
root = merkle_root(leaves)
proof = inclusion_proof(4, leaves)
print("root :", root.hex())
print("valid:", verify_inclusion(leaves[4], 4, len(leaves), proof, root)) # True
tampered = leaves[4].replace(b"E_SLIP", b"E_NONE")
print("tampered:", verify_inclusion(tampered, 4, len(leaves), proof, root)) # False
In production, hash the exact bytes you store, such as the serialized ROS message, rather than a re-encoded view, and store intermediate nodes so proofs are generated without rehashing.
Append-only logs and consistency proofs
A single Merkle root protects a fixed batch. Machine data, however, is a stream. The design that solves this is the append-only Merkle log used by Certificate Transparency, specified in RFC 6962 and its successor RFC 9162.
Records are appended as new leaves of one ever-growing tree. Periodically the log publishes a signed tree head: tree size, timestamp and root hash, signed by the log's key. Two kinds of proof then cover the whole lifecycle:
- Inclusion proofs show that a specific record is in the tree of size n.
- Consistency proofs show that the tree of size m is a prefix of the tree of size n: every record that existed at size m is still there, unchanged and in the same position. Like inclusion proofs, they need only O(log n) hashes.
Consistency proofs make the log tamper-evident over time. If someone deletes last month's error record and recomputes the tree, the new root cannot be proven consistent with the signed tree head a customer or auditor saved last month. The remaining attack, showing different histories to different parties, is countered by having several parties store and compare tree heads.
Signing, anchoring and batching on real devices
A root hash proves integrity relative to itself. To prove that a root came from a particular robot at a particular time, you need signatures and time anchoring.
Signing roots with device keys
Give each device its own key pair and sign every tree head. Ed25519 is a good default: 32-byte public keys, 64-byte signatures, deterministic signing and fast verification.
import time
from cryptography.hazmat.primitives.asymmetric.ed25519 import Ed25519PrivateKey
device_key = Ed25519PrivateKey.generate() # in production: inside secure hardware
tree_head = f"amr-07|{len(leaves)}|{int(time.time())}|{root.hex()}".encode()
signature = device_key.sign(tree_head)
# Verifier side: raises InvalidSignature if the tree head was altered.
device_key.public_key().verify(signature, tree_head)
Register each public key in your fleet registry at provisioning time.
Secure elements and TPMs
If the private key sits in a file on the robot's SSD, anyone with root access can sign forged history. Generate and keep it inside a TPM 2.0 or a secure element, where the host can request signatures but never read the key. Note that many TPMs and secure elements support ECDSA on P-256 but not Ed25519; in that case use ECDSA P-256, and the rest of the design is unchanged. Hardware monotonic counters, where available, also help detect rollback of the log to an earlier state.
Time anchoring
Device clocks drift and can be set by an attacker. To prove data existed no later than a given time, anchor tree heads externally:
- Trusted timestamping. Send the root hash to an RFC 3161 timestamp authority, which returns a signed token binding the hash to its clock.
- Public ledger anchoring. Periodically publish an aggregate Merkle root of many devices' tree heads in a public blockchain transaction, provably existing before that block.
- Cross-anchoring. Including the latest server-acknowledged root in each device batch ties device history to server time.
Batching strategy on constrained devices
Signing every record is wasteful. A practical pattern for robots and IoT gateways:
- Hash each record or each file chunk as it is written, appending the leaf to a local append-only tree.
- Every N seconds or M records, for example 10 seconds or 10,000 records, sign a tree head.
- Transmit the signed tree head immediately; it is about 100 bytes and fits through the thinnest uplink. Upload the bulk data later, when bandwidth allows.
- On upload, the server verifies that the data reproduces the committed roots and stores consistency proofs between successive tree heads.
Because the commitment travels first, data altered on disk while waiting for upload is detected. As an illustrative estimate, 1,000 records per second with 10-second batches means about 20,000 SHA-256 operations per batch, negligible on an application processor; on microcontrollers, use the hardware hash accelerator many of them include.
From Merkle trees to content addressing
The same idea scales to files. In IPFS, a file is split into blocks arranged in a Merkle DAG, and its content identifier (CID) is derived from the root hash, so you can verify it whoever serves it. Filecoin adds proofs that storage providers keep holding the data over time.
For machine data this closes the loop: the device commits to records and files in a signed Merkle log; files are archived under their CIDs; and the log's leaves can reference those CIDs. An auditor can trace a video clip from an archive back to the signed tree head the robot published when it recorded it. We discuss the storage side in hybrid cloud storage for robotics.
As the name suggests, this is the foundation MerkleBot is built on. The MerkleBot platform connects machine data from ROS and ROS 2 robots, industrial arms and IoT fleets, hashes it close to the source and stores it in content-addressed, hybrid storage. Verifiability enables data-driven business models such as Smart Lease, where usage data drives lease terms, as when Southie Autonomy Works obtained a Mitsubishi Electric RV-7FRL arm with no upfront CAPEX. It equally matters for a Universal Robots arm streaming 3D product scans to a cloud AI for defect and fraud detection. For more on these models, read robot-as-a-service and data-driven business models.
Threat model and design checklist
Threat model
| Threat | Mitigation |
|---|---|
| Record edited after recording | Leaf hashes in a Merkle tree; inclusion proof fails against the signed root |
| Records deleted or history rewritten | Append-only log with consistency proofs; tree heads stored by multiple parties |
| Forged log from a fake device | Per-device signing keys registered at provisioning; signature verification on every tree head |
| Device key extracted from disk | Key generated and held in a TPM or secure element; device certificate revocation |
| Backdated data or manipulated clock | RFC 3161 timestamps or public ledger anchoring; cross-anchoring with server roots |
| Split view to different parties | Gossip and comparison of signed tree heads between operator, customer and auditors |
| Leaf/node confusion (second preimage) | Domain-separated hashing (0x00 leaves, 0x01 nodes) as in RFC 6962 |
| Archive file swapped in storage | Content addressing; verify CID or file hash on every retrieval |
| Falsified sensor input | Secure boot, signed firmware, sensor plausibility and cross-sensor checks |
Practical design checklist
- Choose and version the hash algorithm; hash stored bytes, not re-encoded views.
- Use domain-separated leaf and node hashing; follow RFC 6962/9162 tree construction rather than inventing one.
- Commit on the device, as close to the sensor as possible.
- Sign tree heads with a per-device key in a TPM or secure element; register public keys at provisioning.
- Batch by time or record count, and send signed tree heads ahead of bulk data.
- Anchor tree heads to an external timestamp source on a fixed schedule.
- Store consistency proofs between successive tree heads, and let customers and partners keep copies.
- Reference archived files by CID or content hash from the log.
- Monitor continuously: alert on failed signature, inclusion or consistency checks, as part of your fleet monitoring.
Frequently asked questions
Do I need a blockchain to build tamper-evident logs?
No. Merkle trees, signatures and consistency proofs provide tamper evidence on their own; Certificate Transparency works without a blockchain. A public ledger is one optional way to anchor roots in time so that no single party controls the timestamp.
How much overhead does a Merkle log add to an IoT device?
Very little: about two hashes per record and one signature per batch, with signed tree heads around 100 bytes. On microcontrollers, use the hardware hash accelerator if available.
What is the difference between a hash chain and a Merkle tree?
A hash chain links each record to the previous one, so verifying a single record means walking the chain, which is O(n). A Merkle tree supports inclusion and consistency proofs in O(log n), which matters when logs contain millions of records.
Does tamper-evident logging encrypt my data?
No. Integrity and confidentiality are separate. A Merkle log proves data has not changed but does not hide it. Encrypt sensitive data separately, and decide whether leaves hash plaintext or ciphertext based on who needs to verify.
Conclusion
As robot data becomes the basis for payments, warranties, insurance and AI models, data integrity in IoT and robotics moves from nice-to-have to requirement. Merkle trees give you compact commitments and logarithmic proofs; append-only logs with consistency proofs make history verifiable; device keys in secure hardware and time anchoring make it attributable and dated. None of it is exotic, and all of it is cheaper to design in from the start than to retrofit after the first dispute.
If you want verifiable data across your fleet without building the infrastructure yourself, book a 30-minute demo and we will show you how MerkleBot makes machine data verifiable from device to archive.





