We present a universal approach to uncover and correct systematic local errors in complex speech-to-text systems. Whereas previous work to minimize speech recognition errors mostly relies on N-best lists or word lattices, our approach is merely based on the first-best system output. The paradigm of Transformation-Based Learning is adapted from tagging-like applications to the more complicated task of text transformation which obstructs several basic TBL steps. On a professional spontaneous dictation task (including postprocessing and text formatting) we achieve error reductions of 9.6%rel on held-out test data. A special benefit of the approach is the easy interpretation of the learned rules which may serve for diagnostic purposes.