Images and videos are an integral part of any social network. However, these social media platforms can also prove to be the usual suspects for spreading fake news and hate speech. To minimize the impact of such negative incidents, companies like Facebook use automated tools.
In a recent post , Facebook described a new system based on machine learning that can recognize whether text is present in videos and images and transcribe it into a machine-readable format.
This system, called Rosetta, goes a step beyond traditional optical character recognition (OCR) techniques, as it attempts to understand the context of the image and text together.
The company used a billion images and videos from Facebook and Instagram to train the text recognition model. The overall process involves two steps: detecting a rectangular area that may contain text, and then performing text recognition using a convolutional neural network (CNN).
The model is not limited to just English text. It supports various languages and encodings, including Hindi and Arabic.
The Rosetta system is already widely integrated into various Instagram and Facebook products. The company says it also uses it to improve photo search, display personalized content, and identify hate speech content in various languages.
