AIP v1 deterministic encoding
Canonical objects use the core deterministic encoding requirements in RFC 8949 section 4.2.1. V1 restricts CBOR to a small explicit data model so independent encoders do not need to agree on floating-point or tag policies.
Allowed values
- Unsigned integers, major type 0, range 0 through 2^64-1. Used for envelope versions and executed-evidence durations, limits, and exit codes.
- Byte strings, major type 2. Artifact content is a byte string, not UTF-8 text.
- UTF-8 text strings, major type 3. Invalid UTF-8 MUST be rejected. No Unicode normalization is performed. Different Unicode sequences are distinct values.
- Arrays, major type 4, with schema-defined element types and ordering.
- Maps, major type 5, with text keys only.
false(f4),true(f5), andnull(f6). Null represents missing sides of operation effects and unavailable exit codes in executed evidence.
Negative integers, floating-point values, tags, undefined, other simple values, and indefinite lengths MUST be rejected. Empty data is still encoded explicitly: empty array 80, empty map a0, empty bytes 40, empty text 60.
Encoding rules
All lengths and integer arguments MUST use their shortest encoding: inline for 0..23; additional byte for 24..255; 16 bits for 256..65535; 32 bits through 4294967295; 64 bits otherwise. Multi-byte arguments are unsigned big-endian.
Map keys MUST be sorted by unsigned lexicographic comparison of their complete CBOR-encoded bytes, not insertion order or Go field order. With text-only keys, this is equivalent to sorting by UTF-8 byte length then lexicographic UTF-8 bytes. Duplicate keys MUST be rejected, not overwritten. Arrays of IDs MUST be sorted by lowercase hex text with duplicates removed by constructors; readers MUST reject duplicates or unsorted arrays. Other arrays retain their schema order. Effects are sorted by path’s UTF-8 byte sequence, not CBOR map-key order.
The exact envelope is a three-key map: type, version, data. Its encoded key
order is data, type, version. Every data field from objects.md is required.
Version is scoped to the type. Evidence v2 uses this same encoding profile;
the other types and reported evidence retain version 1.
No aliases, omitted defaults, or extra fields are permitted. IDs are 64-byte
lowercase ASCII hex text, not raw digest bytes. Times use the exact textual
format in objects.md, never CBOR date tags. Identity excludes filesystem times,
compression, presentation JSON, local references, and derived reverse links.
An implementation MUST consume exactly one complete object with no trailing bytes. It MUST reject noncanonical encodings even if decoding gives the same logical value. Decode then re-encode and compare is a valid implementation.
Identity and limits
ObjectID = lowercase_hex(SHA-256(encoded_envelope)).
There is no additional prefix or framing in the hash input or stored file. Type and version are already covered by the hash. Each object file contains exactly those bytes. Identical logical objects under these schemas MUST produce identical bytes and IDs, independent of map construction order or implementation. Timestamps are logical data; newly recorded intents/evidence can have different IDs even when their other fields agree. States and operations have no timestamps.
The v1 local profile limits a complete encoded object to 67,108,864 bytes and CBOR nesting depth to 64 below its root value. Lengths MUST be checked against remaining bytes before allocation; implementations SHOULD allocate collections incrementally. The graph verifier limits a reference path to 4096 edges; deeper graphs return a resource-limit error. This is an implementation limit, not a claim that a larger graph is cryptographically invalid.
Golden vectors cover every object type, the empty root, binary bytes, null effects, references, timestamps, and exact hashes. Codec tests additionally exercise length boundaries, unordered and duplicate map keys, overlong lengths, invalid UTF-8, truncation, unsupported types, and excess nesting.