
The term "Generative AI" refers to AI models and algorithms that can generate new content or data similar to the data they were trained on. This includes a wide range of different content types across many disciplines:
- Text (articles, product descriptions, letters etc.)
- Images (logos, designs, photograph adjustments etc.)
- Sound (sound effects, theme tunes etc.)
- Video (short animations, instruction videos etc.)
- Science (molecular structures, drug discoveries etc.)
- Interaction (customer support, companionship etc.)
How Generative AI Works
The neural network of Generative AI models consists of multiple layers which vary based on the type of data to be generated. This is fed with large amounts of training data to teach the model how to generate something similar, with adjustments to the weights and parameters of it's neurons made to minimize errors between the generated data and the training data. Once training is complete, new data can be generated with a high degree of accuracy from a starting sequence or value (usually called a "prompt"). It is important to continue to fine-tune the model by training it with new data and evaluating the quality and relevance of it's output. Different Generative AI models differ in the type and extent of user involvement required during learning, with some requiring active involvement and others learning completely unsupervised.
Today, Generative AI technology mostly involves specialized neural networks called "transformer models" but there are many different kinds of neural networks involved:
- Generative Adversarial Networks (GANs): Consist of a generator and a discriminator and are often used to create realistic images.
- Recurrent Neural Networks (RNNs): Specifically designed for processing sequential data like text and are used for generating text or music.
- Transformer-based models: Used for text generation. ChatGPT (Generative Pretrained Transformer) from OpenAI is a popular example.
- Flow-based models: Used in advanced applications to generate images or other data.
- Variational Autoencoders (VAEs): VAEs are frequently used in image and text generation.
- Diffusion models: Generate data by progressively removing noise from a random input and are mainly used in realistic image generation. Stable Diffusion is a popular example.
Example Generative AI Models
Here are some current examples of the most popular Generative AI models:
- ChatGPT: This text generator is an AI chatbot powered by OpenAI's GPT-4 language prediction model. Because it considers the user's conversation history and is trained on large amounts of text data, ChatGPT simulates a more natural style of conversation.
- DALL-E: This image generator was trained on a large amount of images and associated text descriptions, meaning it can connect the meaning of words with the visual equivalent.
- Gemini: This text generator is Google's AI chatbot powered by the Large Language Model Gemini 1.5 and draws it's data from the internet.
- Claude: This text generator is Anthropic's AI chatbot, and is founded by former employees of OpenAI. Claude is also an extremely popular AI assistant in the scientific and coding communities.
- LLaMA: This is the latest model from Meta. Various versions are freely available and well-suited for custom AI applications, making it very appealing to those who want to avoid proprietary providers.
Potential Problems of Generative AI
The quality and accuracy of the output of AI models is limited by the quality and accuracy of the data used to train them, as well as the amount of human adjustment and correction applied. For example, it is hard to identify misleading information or measure any bias in the original data source, and so outputs should always be checked for plausibility and quality. And even when the output is exactly what the user asked for, it can be used for nefarious purposes, such as making public figures appear to say or do things they didn't, or using the technology to hack into private computer networks.