Microsoft Germany announced GPT-4 as a new multimodal AI model at an AI kickoff event. According to Microsoft Germany Chief Technology Officer Andreas Braun, the model is to be officially presented next week.
Successor to the ChatGPT-LLM
< p class="p text-width">GPT-4 was a multimodal Large Language Model developed by OpenAI. In the context of AI models, multimodality refers to the ability to deal with several types (modalities) of inputs and outputs – for example, video and audio in addition to text. GPT stands for “Generative Pre-trained Transformer” – a type of artificial neural network that has been trained on a large data set in order to be able to generate outputs. The previous GPT models – which were recently used in the OpenAI chatbot ChatGPT – worked exclusively on text. In addition to the previous GPT models, research partner OpenAI has gained experience with generative models in other modalities – for example DALL-E for text-to-image and Whisper for audio-to-text generation.
Interested readers can learn more details about how language models work in the article “Large Language Models and Chatbots: The Technology Behind ChatGPT, Bing Chat and Google Bard”.
Few details about the model
At the event, Microsoft gave few details about the upcoming model. In principle, artificial intelligence was advertised as a “game changer” that could make certain repetitive tasks significantly easier. Senior AI Specialist Clemens Sieber gave an example of a Dutch call center with 30,000 calls per day that saved up to 500 working hours per day by automatically transcribing phone calls to text. Regarding the problem of Large Language Models inventing facts without supporting evidence, Siebert stated that Microsoft is working on a feedback loop. A similar process is already being used at ChatGPT, where users can rate outputs of the model either with a thumbs up or down.
CB-Funk Podcast Episode #10: “More appearances than Being” also applies far too often to hardware