The first piece of real logic written for Ferro was the error type. Not the tensor, not any arithmetic — the thing that explains what went wrong. That ordering is deliberate, and the reasoning is that in a numerical library the error messages are not diagnostics you occasionally consult. They are the primary interface you interact with while building anything.
#The shape of the problem
Almost every operation in a tensor library has a compatibility requirement. You can add two arrays of the same dimensions. You can multiply a 3×4 matrix by a 4×5 matrix, but not by a 5×4 one. You can sum along dimension 2 of a three-dimensional array, but not of a two-dimensional one.
These requirements are checked at runtime, because the dimensions are runtime values. And they are violated constantly — not because people are careless, but because keeping track of the dimensions of a dozen intermediate values through a neural network is genuinely hard. Getting a shape wrong is the single most common mistake in the entire field.
So the question is what happens when one is violated. The lazy answer, which most code reaches for, is an assertion:
assert_eq!(lhs.shape(), rhs.shape());That produces something like:
thread 'main' panicked at src/ops/binary.rs:47:5:
assertion `left == right` failed
left: [2, 3, 4]
right: [2, 3, 5]Which is not terrible. It has both shapes. But it does not say what operation
failed — binary.rs:47 is the shared implementation for a dozen operations —
and it does not say which of the two operands was yours and which came from
somewhere else in the pipeline. In a stack twenty calls deep inside a training
loop, "some binary operation at line 47 got mismatched shapes" leaves you
reading code to find out which one.
#What Ferro does instead
Every shape-related failure carries the name of the operation and both shapes:
shape mismatch in `add`: [2, 3, 4] vs [2, 3, 5]You know the operation. You know both shapes. You can see immediately that the last dimension is the problem and that it is off by one, which usually points straight at whatever produced the second operand.
The same treatment applies across the board. Reshaping computes and reports element counts, because "24 into 25" tells you instantly that you are off by a row, in a way that "(2,3,4) into (5,5)" does not:
cannot reshape [2, 3, 4] (24 elements) into [5, 5] (25 elements) in `reshape`Dimension indices report the rank, because the mistake is nearly always counting from the wrong end:
dimension 4 out of range in `squeeze`: shape [2, 3] has rank 2And a few include the remedy directly, when there is one obvious fix:
`gemm` requires a contiguous tensor; call `.contiguous()` first
device mismatch in `add`: Cpu vs Cuda(0) (move one with `.to(device)`)None of this is clever. It is just a decision to spend a sentence on each error variant, made once, before there were a hundred of them to retrofit.
#How Rust handles errors
Worth a short detour, because Rust's approach differs from most languages.
There are no exceptions. A function that can fail returns a Result, which is
a value that is either success or failure, and the compiler will not let you
use the contents without acknowledging that it might be the failure case. This
is more verbose than exceptions and considerably harder to ignore — you cannot
accidentally let an error propagate silently, because "silently" is not
available.
The convention is one error type per library, an enum listing every way things
can go wrong. Ferro's lives in ferro-core and is shared across every crate,
so ferro-nn and ferro-rl all report failures in the same vocabulary rather
than each inventing their own.
The library that makes this pleasant is thiserror, which generates the
message-formatting code from an annotation:
#[error("shape mismatch in `{op}`: {lhs:?} vs {rhs:?}")]
ShapeMismatch {
op: &'static str,
lhs: Vec<usize>,
rhs: Vec<usize>,
},The #[error(...)] line is the message template, and the fields are the data.
You write the message next to the thing it describes, which is a small detail
that matters a lot for whether the messages stay accurate as the code changes.
#Two decisions in that snippet
op is &'static str, not String. In Rust, a String owns
heap-allocated memory, while &'static str is a reference to text baked into
the compiled binary. Since operation names are always literals like "add",
the second works and costs nothing.
This means constructing an error never allocates memory. That sounds like
premature optimization for something that only happens when things break —
except that Ferro will eventually run on a memory pool where allocation is
carefully managed, and the gradient checker planned for later deliberately
triggers thousands of failures in a loop while probing edge cases. An error
path that allocates is a small nuisance in both. Making it not allocate cost
nothing here and would be tedious to change later, once every construction
site in the workspace passes a String.
The shapes are Vec<usize>, formatted with {:?}. Rust's debug formatting
of a list of numbers is [2, 3, 4] — exactly the notation everyone already
uses for array shapes, and identical to how NumPy and PyTorch print them. No
custom formatting code required.
#The type is marked non-exhaustive
One line above the enum:
#[non_exhaustive]
pub enum Error { ... }This tells code outside the crate that it must not assume the list of variants
is complete — any match on a Ferro error needs a catch-all arm.
The reason is that the error type starts with eleven variants and the finished library will have many more. Without this annotation, adding a variant is a breaking change: every place anyone matched on the error stops compiling. With it, new variants can be added freely.
#Testing the message text
The error type ships with three tests, and they assert on the strings:
#[test]
fn shape_mismatch_message_carries_both_shapes_and_op() {
let err = Error::ShapeMismatch {
op: "add",
lhs: vec![2, 3, 4],
rhs: vec![2, 3, 5],
};
let msg = err.to_string();
assert!(msg.contains("add"), "message must name the op: {msg}");
assert!(msg.contains("[2, 3, 4]"), "message must show lhs: {msg}");
assert!(msg.contains("[2, 3, 5]"), "message must show rhs: {msg}");
}Testing the exact wording of an error message is often a bad idea — it makes tests brittle against harmless rewording. Here it is the right call, because the content is the feature. The requirement is not "produces some message", it is "names the operation and shows both shapes". A future refactor that drops one of the shapes to simplify a signature would be a regression, and this test is what catches it. The assertions check for the presence of the three facts, not the sentence structure around them, so rewording stays free.
#Why any of this before the tensor exists
The honest reason: because it will not happen otherwise.
Error handling written after the fact is error handling written under pressure,
while you are trying to finish something else. It gets the minimum. Someone
adds assert!(a.shape() == b.shape()) because it takes four seconds and the
real work is elsewhere, and that is completely reasonable in the moment. Then
it happens forty more times, and now improving the errors is a forty-site
refactor that never reaches the top of anyone's list.
The next post is about a related piece of infrastructure with the same logic: the automation, built before there is anything much to automate.