HomeScience & TechnologyGoogle uses MLPerf to demonstrate performance on giant...

Google uses MLPerf to demonstrate performance on a giant version of BERT


Google is leveraging MLPerf competitors to demonstrate general performance on the giant variant of its BERT language product (Bidirectional Encoder Representations from Transformers is a transformer-based machine learning technique for pretraining natural language processing developed by Google).

The deep learning world of artificial intelligence is obsessed with size.

Deep learning programs, such as OpenAI's GPT-3, continue to use more and more GPU chips from Nvidia and AMD to build ever larger software programs. The accuracy of the programs increases with size, researchers say.

See also: easiersearchesyour

Google BERT MLPerf

This obsession with size was on full display Wednesday in the latest industry benchmark results reported by MLCommons (a global and open nonprofit dedicated to improving machine learning for everyone), which sets the standard for measuring how quickly computer chips can crack the deep learning code.

Google decided not to submit to any of the standard deep learning benchmark tests, which consist of programs that are established in the field but are relatively outdated. Instead, Google engineers presented a version of Google's BERT natural language program that no other vendor used.

MLPerf, the benchmark suite used to measure performance in competition, reports results for two segments: the standard “Closed” segment, where most vendors compete on established networks like ResNet-50, and the “Open” segment, which allows vendors to test non-standard approaches.

See also: Google: Cryptocurrency miners are hacking Cloud accounts

In the Open section, Google showed off a computer using 2,048 Google TPU chips, version 4. This machine was able to develop the BERT program in about 19 hours.

The BERT program, a neural network with 481 billion parameters, has not been previously disclosed. It is more than three orders of magnitude larger than the standard version of BERT that is currently in circulation, known as “BERT Large,” which has just 340 million parameters. Having many more parameters usually requires much more computing power.

MLPerf

Google said the new submission reflects the growing importance of scale in artificial intelligence.

The MLPerf test suite is a creation of MLCommons, an industry consortium that issues multiple annual computer benchmarking assessments for the two parts of machine learning: so-called training, where a neural network is created by improving its settings over multiple experiments, and so-called inference, where the completed neural network makes predictions as it receives new data.

Wednesday's report is the final benchmark test for the training phase. It follows the previous report in June.

The full MLPerf results were discussed in a press release on the MLCommons website. Full details of the submissions can be seen in tables posted on the website.

Google BERT MLPerf

Google's Selvan said MLCommons should consider including more large models. Older, smaller networks like ResNet-50 "only give us a proxy" for large-scale training performance, he said.

The missing piece, Selvan said, is the so-called efficiency of systems as they grow. Google managed to run its giant BERT model with 63% efficiency, he told ZDNet, as measured by the number of floating-point operations per second performed relative to a theoretical capacity. That, he said, was better than the next highest industry result, 52%, achieved by Nvidia when developing the Megatron-Turing language model it developed with Microsoft.

David Kanter, executive director of MLCommons, said that the decision on large models should be left to the Commons members to decide collectively. He noted, however, that using small neural networks as testbeds makes the competition accessible to more places.

See also: How to upload files and folders to Google Drive?

In contrast, MLPerf's standard tests, whose code is available to everyone, are a resource that any AI researcher can tap into to replicate the tests, Kanter said.

Google has no plans to release the new BERT model, Selvan told ZDNet in an email, describing it as “something we did just for MLPerf.” The program is similar to designs described in Google research earlier this year on highly parallelized neural networks, Selvan said.

Despite the novelty of the 481 billion parameter BERT program, it is representative of real-world tasks because it is based on real-world code.

Information source: zdnet.com

📧
Subscribe to the SecNews Newsletter

The most important Security & Technology news in your Inbox.

Teo Ehc
Teo Ehchttps://www.secnews.gr
Be the limited edition.

SEARCH

FOLLOW US

📧
Newsletter SecNews
The most important Security & Technology news in your inbox.

LIVE NEWS