In June we wrote about passing 10,000 automated tests, the small programs that check NomNom still does what it should every time we change anything. We finished that post with “here’s to the next ten thousand.” It took about three months to add the next twenty thousand.
This morning, while writing this, we ran everything: 33,398 tests passed, and none failed.
We said in June that a test count on its own proves very little, and that is still true. So this post is not really about the number. It is about the two things that changed, and about what the tests still missed along the way.
The app is tested now, not just the server
The June post was about the server: the part of NomNom that takes bookings, charges deposits and sends emails. The app your team actually uses — on the iPad at the host stand, on a phone, or in NomNom Web — had far less coverage.
It now has 7,405 tests of its own. They check what the tests on the server cannot: that a booking appears in the right place on the chart, that pressing Save sends exactly what you typed, that a dialog fits on the screen, and that a button does what its label says. The two sets run separately, and a change to either one is not finished until its tests pass.
What 33,000 tests still didn’t catch
Here is the honest part. This summer, four problems reached real restaurants while every test was passing. Each one taught us something about how to test, not just how much.
The card that only broke in the finished app
A new “Get started” card looked perfect while we were building it and in every test. On a venue’s screen, its text came out one letter per line.
The fault only showed in the finished version of the app, not in the version we test with. Our tests were drawing the screen without NomNom’s real styling, and the fault lived in that styling. The tests checked that nothing crashed; nothing did. The words just got crushed to nothing.
What changed: screen tests now load the app’s real styling and measure the result: how wide the text is, how tall. “It didn’t crash” is not the same as “it looks right.”
The text that passed every test and still couldn’t be read
On the evening a venue went live on NomNom Web, their staff told us they were struggling to read the lighter grey text. They were right. Our secondary text colour had a contrast of about 2.6 to 1 against a white background, where the accessibility guidelines ask for at least 4.5 to 1, and it was used for most of the hints, labels and times on the screen.
Every test passed, because every test checked that the text was there. None checked that a person could read it.
What changed: we darkened the colours, and a test now works out the contrast of every text colour against every background it sits on, using the formula from the web accessibility guidelines. If anyone makes the text paler again, the test fails.
Tests that read the code instead of running it
A restaurant set up an event, and guests booking it online were skipped past a step they needed to see. The online booking page had plenty of tests. The trouble was what they checked: that each piece of the page’s code contained the right instructions. None of them clicked through the page the way a guest does, one step after another, and the bug only happened between two steps.
What changed: when a problem spans more than one step, we now run the real booking page in a real browser, using the venue’s real availability, and reproduce the bug before we fix it. It is not called fixed until the same run shows it gone.
The change we didn’t write
NomNom’s server is built on an open-source framework, and in September we took an update to it. It was a security improvement, and we changed none of our own code. For one afternoon, staff looking for a free table in the app were told there were none, until we rolled the update back by hand.
The update had started rejecting web addresses that contain a colon. Exactly one of ours does: the one the app uses to ask which tables are free at a particular time. Our tests checked that part of the server directly and skipped the framework’s front door, so they never saw the rejection.
What changed: a test now sends that exact address through the framework’s real rules, so the next update that rejects it fails our tests instead of reaching your staff.
What we learnt about testing
Beyond those four, a pattern kept coming up this summer: a test that passes is not proof that anything is protected. A few times we found tests that were checking a copy of the thing that mattered. One was a helper with a full set of tests, all passing, that nothing in NomNom ever called. The tests were green and the feature did not exist.
So a few rules got stricter:
- Every fix is proven twice. We put the bug back, watch the new test fail, then restore the fix and watch it pass. A test that would have passed with the bug still in place proves nothing, so it doesn’t count.
- Test what actually runs. The test has to go through the same path a real booking does, not a helper beside it.
- Check every one, including the next one. Where we can, a test checks a rule across every email, page or colour at once, rather than a list someone wrote down, so the next one anybody adds is covered automatically.
- The rule covers the app too. “No change without a test” was first written down for the server, and an app fix once went out without one because of that gap. The rule now covers both, and an automated check flags any change that arrives without a test.
If you’re switching booking system this month
If you’re one of the restaurants moving on from Quandoo, this is the part that matters to you. The fear with any new system is that it breaks on a Saturday night. We can’t promise nothing will ever go wrong; this post is four examples of things that did. What we can tell you is that every one of them was found, fixed, and then locked out by a test so it cannot quietly come back. That happens every day, on every change, before it reaches your restaurant.
For the developers reading this
- Server: 25,993 DUnitX tests across 527 test units, Delphi / DMVCFramework. The full run takes about 20 seconds, and still never touches a real database: everything runs against in-memory fakes.
- App: 7,405 Flutter unit and widget tests across 500 test files, in just over three minutes. 18 more are skipped on purpose; most of them only apply to a white-label build of the app.
- Mutation-proven fixes. Every bug-fix test is run against the unfixed code first, and must go red.
- Layout tests assert geometry, under the production theme, not just the absence of an exception. A release build has no layout asserts to save you.
- Accessibility is a unit test: the WCAG relative-luminance formula in plain Dart, checked first against the spec’s own worked values, then asserting 4.5:1 for every text tier on every surface.
- Sweeps end with a floor. A test that walks every file must also assert that it saw at least N of them. Otherwise, a renamed folder turns it into a test of nothing that passes forever.
We’ll write again when there is something worth saying, whether that’s another milestone or another thing the tests taught us. Probably the second.