This paper proposes to use combined acoustic and visual feature vectors to distinguish live synchronous audio-video recordings from replay attacks that use audio with a still photo. Equal error rates below 2 % are achieved using a multi-dimensional eigenlip representation and EERs of 7% are achieved with a one-dimensional lip-opening ratio.