How to tell if a productivity change actually worked

· 6 min read

To tell whether a change worked, you need a baseline from before the change, one metric chosen in advance, and at least two weeks after. Most people skip the baseline, which is why every change feels like it worked for a week and then nothing is different. Per-app time on your Mac gives you a metric you did not have to produce and cannot easily flatter.

Why every change feels like it worked

The first week of any productivity change is a good week. You moved the standup, you closed chat until 11, you bought the standing desk. You are paying attention to your day in a way you were not before, and attention on its own improves things for a few days. Then attention fades, the day returns to its shape, and you are left with a memory of a good week and no way to tell whether the change did anything.

The memory is the problem. Feelings about a week are made on Friday afternoon from whatever happened last. A week with one bad Thursday feels like a bad week. Nothing in that tells you whether the editor got more of your day.

What you need is a number recorded while the week happened, chosen before the change, and comparable with the same number from before. That is a measurement. Everything else is a mood.

Choose one metric before you start

The metric has to be chosen first, because if you choose it afterward you will pick whichever one moved.

Good metrics for a personal productivity change are things a per-app record can show directly:

  • Daily time in your main work app (editor, design tool, writing app). This is the usual one.
  • Daily time in chat and email combined. For changes aimed at interruptions.
  • The day total. For changes aimed at working less, or at least not more.
  • The number of days the closing time held. For changes aimed at the end of the day.

Pick one. Write it down with the change and the date: “From Monday, chat closed until 11. Metric: editor time per day.” A second metric is allowed only as a check that the first did not improve by wrecking something else; editor time up while the day total is up by the same amount is a longer day, not a win.

Do not choose “how focused I felt”. That is the thing being tested.

The before-and-after protocol

  1. Record a baseline. Two normal weeks before the change, with the metric written down every day. If you cannot wait two weeks, take one, and accept that the comparison will be rougher.
  2. Write down your prediction. “Editor time will go from about two hours to about three.” A prediction makes the result mean something; without it you will look at the number and decide afterward whether it counts.
  3. Make the change on a Monday and change nothing else. If you also start a new project that week, you will not know which one moved the number.
  4. Keep recording for two weeks after. Do not look at the running total during the first week if you can help it; the week-one glow is exactly what you are trying to see past.
  5. On the second Friday, line up the four weeks. Baseline average against after average, for the one metric.
  6. Compare with the prediction. Three outcomes: it moved as predicted, it moved less than predicted, it did not move.

Two weeks after is the minimum because the first week is contaminated by attention. If the second week holds the gain, the change probably worked. If the second week drops back to baseline, it was the attention, not the change.

Reading the result honestly

Daily numbers bounce. A single day can be an hour off the average for reasons that have nothing to do with your change: a long meeting, a sick child, a release. So compare week averages, not days.

A useful check: look at the two baseline weeks against each other. If they differ by forty minutes a day with no change at all, then a forty-minute improvement after the change is noise. You need the after weeks to sit clearly outside the range the baseline weeks showed.

Watch for the change that moves the metric the wrong way through a side door. Closing chat until 11 can raise editor time and also raise the day total, because the chat got done after 6 instead. That is the second metric’s job to catch. Are you working more from home or just longer? is about that specific failure.

And be willing to get “it did not move”. That is a successful experiment. You have learned that the thing everyone recommends does nothing for your week, and you can stop doing it.

Keeping the record without effort

The protocol asks for a number every day for four weeks. That is twenty entries, and the method dies if any week goes unrecorded. The record has to be made without you.

Punchcard is built for this. It sits in the menu bar and notices which app is in front during the day, by app name only, with no macOS permissions. At the closing time you set it prints a receipt: one line per app with the time, a day total, a stamp. Your metric is one line on it, every day, whether you remember the experiment or not. The week roll on Sunday gives you the weekly average for step 5 without adding anything up, and a CSV export lets you put all four weeks in one sheet if you want to look closely.

It tracks from the day you install it, with no account, so the baseline can start today. The first seven receipts are free; after that it is a one-time purchase.

Be clear about the limits. It shows time per app, not what happened in the app, so “editor time” means the editor was in front, which is a ceiling on work and not a measure of it. It cannot tag weeks as “before” and “after”; you keep that note yourself. And it cannot measure anything that does not show up as an app in front.

Changes that cannot be measured this way

Some changes do not show up in per-app time, and it is better to know that before the experiment than after.

Quality changes are invisible. A change that makes your code better or your writing clearer will not move the editor line. The only evidence for those is the work itself.

Changes to what happens away from the Mac are invisible. Walking meetings, paper sketching, thinking in the shower. If the change is “think more before typing”, the editor line may go down and the change may still have worked.

And changes to how you feel are real even when they are not measurable. If closing chat until 11 did not move the editor line but made the mornings calmer, that is a result. It is just not one this method can show you, and it is worth being honest that you are keeping the change for a reason the numbers do not support.

Questions

How long should the baseline be?

Two weeks if you can, one if you cannot. The point of the baseline is to learn how much your days bounce with no change at all, so that you can tell a real effect from an ordinary good week.

What if I cannot stop making other changes?

Then you will not know which one worked, and you should say so rather than crediting the one you liked. If you have to stack changes, test the biggest one first on its own and add the others later.

Can I measure a change to a team’s productivity this way?

Not with a personal tracker. Punchcard records one person’s Mac, for that person, with no team view and no sharing. A team measurement needs everyone’s agreement and a different tool.

The number moved but I am not sure it is real. What now?

Run it for two more weeks. If the after weeks hold while the baseline weeks looked nothing like them, it is real. If it drifts back, you have learned something true anyway.