Skip to content

Characterization Tests: Recording Behavior, Not Judging It

Adityo Guni Waluyo

Characterization tests lock in what code actually does while product decisions stay pending, guarded by a pending marker above the assertion.

TL;DR

Characterization tests document current behavior, not correctness; a failure means something changed, not broke. The author found pending-decision markers above such tests and realized the odd procedure—write a failing assertion, replace it with actual output—is intentional. A passing test guards deferred decisions better than any comment, letting owners decide later through one honest commit.

Today I re-ran the test suite for two service modules. Hundreds of green lines scrolled past, then my eyes stopped at one odd spot: right above an assertion sat a literal comment, // characterization: current behavior, owner decision pending. The test passed, but the behavior it locked felt wrong. One of these tests asserts that the per-category report is case-sensitive. That is not a decision, just a note: this is how it behaves today, and someone has to decide later.

My wrong guess about the word "test"

I used to assume a characterization test was a test that confirms the code is correct. It's called a test, so surely it verifies.

Wrong. Michael Feathers' definition, recorded by Alberto Savoia: characterization tests "don't check what the code is supposed to do, as specification tests do, but what the code actually and currently does" [1]. The focus is not truth, it is documentation. Feathers states the purpose plainly: to "document your system's actual behavior, not check for the behavior you wish your system had" [2].

The difference is not just philosophical. With a regular test, a failing assertion means something is broken. With a characterization test, a failing assertion means something changed. Full stop. Whether that change is good or bad is a separate decision that is not the test's job.

The procedure deliberately uses failure

The procedure Feathers describes sounds like a joke: write an assertion you know will fail, run it, "let the failure tell you what the actual behavior is", replace the expectation with the actual value, repeat [1]. Why not just write the correct value up front? Because we often don't know it. The "correct" value is exactly what's in question. The first failure is a measuring instrument, not an accident.

In the suite I read earlier today, the pattern shows up in several places: category-name validation with per-character bounds, notification length counted in raw bytes without trimming, note columns that are never cleared across workflow stages, public ticket subjects that appear without the category prefix. Four behaviors, four pending markers, one repeated sentence: the decision waits for the product owner.

The locked behavior is not necessarily right. Feathers again: "if you haven't determined that the behavior you've uncovered is a bug, it's often a good idea to leave the test in place" [2]. Fixing a bug is a separate commit that openly admits it changes one assertion. What is forbidden is changing behavior quietly through a refactor.

Why deferring a decision needs a test, not just a comment

A comment saying "ask the owner first" stops no one. A passing test in CI stops everyone. While that marker stays pending, the test works as a fence: whoever changes the behavior has to open that assertion and learn that a decision has not landed yet.

Martin Fowler calls a suite like this a safety net: with it you can clean up code safely because "the bug detector will go off" when you slip, and that confidence is what makes people willing to change the system [3]. For a codebase with gray areas like mine, the net is not a luxury. It is what lets a decision stay deferred without the project turning into regression whack-a-mole.

One thing I take away from today: deferring a decision is not a sign of an unhealthy project. Unhealthy is a deferred decision whose behavior nobody guards. A characterization test plus a pending marker makes the delay explicit, guarded, and addressable: whenever the owner decides, one assertion waits to be changed, and one commit honestly admits doing it.

Sources

  1. Alberto Savoia, "Working Effectively With Characterization Tests" (artima.com)
  2. Michael Feathers, "Characterization Testing" (michaelfeathers.silvrback.com)
  3. Martin Fowler, "Self Testing Code" (martinfowler.com)

Related articles