Describing an image is especially easy for humans- but this does not hold for computers.

However, this seems to be starting to change, and an indication is the work of Google, who developed a machine learning that is able to automatically generate captions to describe images the first time it "sees" them.
As written by Oriol Vinyals, Alexander Toshev, Samy Bengio and Dumitru Erhan, scientists at the company on Google's research blog, a system of this kind could help people with vision problems in the long term understand images, provide alternative text for images in parts of the world where network connections are poor and make image search on Google easier.
Recent research has resulted in a significant improvement in detection, logging and tagging/ labeling of objects. However, an accurate description of a complex scene requires a deeper representation of what is happening, «capturing» how the various objects relate to each other and then «translating» the set of «conclusions» into natural language.
«Many attempts to build computer-generated natural image descriptions suggest the combination of modern state of the art techniques both in computer vision and in natural language processing, for the formation of a comprehensive image description approach. But what would happen if instead of that we combined recent computer vision and language models into a single, jointly ‘trained’ system, taking an image and immediately generating a sequence of words – readable by humans- to describe it?» the researchers ask.
Source: naftemporiki.gr
