The four-day week experiment, measured instead of guessed

· 6 min read

Trying a four-day week without measuring anything produces a verdict based on how the month felt, and how a month felt is not evidence. Everyone who tries this reports that they got the important things done, because the important things are what you remember.

The question worth answering is narrower and harder: did output hold, and where did the fifth day’s work actually go? That needs numbers from before as well as during.

Measure the baseline first

The most common mistake is starting the experiment and then wondering what to compare against.

Run a normal five-day schedule for two to three weeks, changing nothing, and record it. You need:

  • Total hours worked per week, honestly, including the evening pieces
  • The distribution across kinds of work
  • Whatever output measure fits your work: pieces shipped, clients served, words, tickets closed

Two weeks is a minimum and three is better, because one atypical week will distort a shorter baseline.

The temptation is to skip this and start the interesting part. Resist it; without a baseline the experiment cannot conclude anything.

What to watch during the trial

Run the shorter week for at least a month. A fortnight is not enough, because the first week is unrepresentative in both directions: novelty helps, and the backlog from the transition hurts.

Three numbers matter.

Total hours. The question is whether the week actually shortened. It frequently does not. Hours migrate into the four remaining days, or into the fifth day informally, and a four-day week that is really four long days plus a morning is a different thing from what was intended.

The shape of the four days. If they became longer and denser, that is a real change with its own costs, and worth knowing rather than discovering later.

Output. Against the baseline. This is the number the experiment exists to answer, and it is the one most easily replaced by an impression.

The finding people usually get

Two things tend to show up, and neither is the one people expect.

The first is that a meaningful part of the fifth day was not productive work. Meetings that existed because there was time for them, correspondence that expands to fill available hours, administration that could be batched. Removing a day forces prioritisation, and some of what gets dropped turns out not to have mattered.

The second is less comfortable: some of the fifth day’s work reappears in the evenings and at weekends, informally and unrecorded. A record catches this and memory does not, because informal work is exactly the kind that does not get remembered as work. If the total hours have not fallen, the experiment has changed the schedule rather than the workload, which is a legitimate outcome but not the one being claimed.

Getting an honest total

The reason this needs an automatic record rather than a log is that the pieces which move are small and scattered. Twenty minutes on a Sunday evening, half an hour before the school run, a call taken on the off day.

Reconstructing a week omits nearly all of that, and it is precisely the part that determines whether the experiment worked. A record that runs without being started catches it.

Why your estimate of your own workday is wrong by hours covers the general problem, and one week of time tracking as an experiment covers running a shorter version of this.

Where Punchcard fits

Punchcard records the frontmost application by name, subtracts idle time, and prints a receipt of the day at the closing time you set. A week of receipts is a week of totals, generated without you deciding anything.

For an experiment like this, the properties that matter are that it runs on the days you did not plan to work, and that it required no discipline to maintain. A Saturday with ninety minutes on it produces a receipt showing ninety minutes, which is the data point the experiment needs and the one you would not have logged.

It has no projects or tags, so the split between kinds of work is by application rather than by category, and any finer allocation is a step you do. A receipt for a day off, and why it is worth printing covers the off-day case specifically, and counting the days covers what a month of them shows.

Reading the result

At the end of the month, four comparisons:

  1. Weekly hours, trial against baseline. Did the week actually shorten?
  2. Output, trial against baseline. Did it hold?
  3. Distribution across the four days. Did they become unsustainably dense?
  4. Off-day and weekend hours. Did work leak into them?

The honest outcomes are: it worked, output held and hours fell; it did not, output fell; or it appeared to work while hours simply moved. The third is the most common and the easiest to miss without a record.

Questions

A month seems long. Can I run it for two weeks? You can, and the result will be dominated by transition effects. A month is the minimum for a conclusion you would act on.

My output is not countable. Then choose a proxy and stay consistent: pieces of work completed, clients handled, whatever unit fits. Consistency matters more than the unit being perfect.

What if the result is ambiguous? That is a real outcome and worth respecting. Ambiguous often means the schedule change was neutral and something else is the constraint. Extending the trial by another month sometimes resolves it.

Does this work for employed people? The measurement does. Whether you can run the experiment is a conversation with an employer, and having baseline data makes that conversation considerably easier to have.