Edge AI
Joules per inference as the design metric
The Joule-Bounded Inference Architecture makes energy budget an admission criterion. Sub-15 W and INT8/INT4 figures are design targets, not measurements.
Design targets · not measured silicon results
- Trust boundary
- Supporting logic
- Signed data in flight
Datacentre design does not transfer to a sealed cabinet
Stable power and active cooling are assumptions edge sites do not share. Scaling a throughput-first part down still leaves it outside the thermal and energy envelope.
- Measured epoch
- Flagged for re-measure
- Light = the digest being extended
Bound energy first, then fit the work
Requests admit under an Energy Contract naming joule budget, model tier and clock profile. INT8/INT4 datapaths and a traffic-aware memory hierarchy target energy per inference; model versions stay measured and signed.
What follows from the diagram
-
Energy Contract admission
Budget is an entry condition, not a post-hoc observation
-
INT8 / INT4 datapaths
Reduced precision cuts arithmetic and memory energy
-
Local dataflow
Weights and activations stay near the arithmetic
-
Measured model version
Results attribute to a specific signed model
Reviewing this mechanism?
The specification can still change. That stops being true after tape-out.