This paper develops a methodology for the design of audio-video data corpora of the speaking face. Existing corpora are surveyed and the principles of data specification, data description and statistical representation are analysed both from an application-driven and from a scientifically motivated perspective. Furthermore, the possibility of "opportunistic" design of speaking-face data corpora is considered.