The top big language model talents only care about these 10 challenges

Source: Silicon Rabbit Racing

Author: Lin Ju Editor: Man Manzhou

Image source: Generated by Unbounded AI

**Editor's note: This article explores the top ten challenges in large language model (LLM) research. The author is Chip Huyen, who graduated from Stanford University and is now the founder of Claypot AI, a real-time machine learning platform. She was previously at NVIDIA , Snorkel AI, Netflix, and Primer develop machine learning tools. **

I am witnessing an unprecedented situation: so many of the world's top minds are now devoted to the unified goal of "making language models (LLMs) better."

After talking to many colleagues in industry and academia, I tried to summarize ten major research directions that are booming:

1. Reduce and measure hallucinations (Editor’s note: hallucinations, hallucinations of AI, that is, incorrect or meaningless parts of AI output, although such output is syntactically reasonable)

2. Optimize context length and context construction

3. Integrate other data modes

4. Increase the speed and reduce costs of LLMs

5. Design a new model architecture

6. Develop GPU alternatives

7. Improve agent availability

8. Improved ability to learn from human preferences

9. Improve the efficiency of the chat interface

10. Building LLMs for non-English languages

Among them, the first two directions, namely reducing "illusions" and "contextual learning", may be the most popular directions at the moment. Personally, I am most interested in items 3 (multimodality), 5 (new architecture), and 6 (GPU alternatives).

01 Reduce and measure illusions

It refers to the phenomenon that occurs when an AI model makes up false content.

Illusion is an inescapable quality in many situations that require creativity. However, for most other application scenarios, it is a drawback.

I recently participated in a discussion group about LLM and spoke with people from companies such as Dropbox, Langchain, Elastics, and Anthropic, and they believe that large-scale enterprise adoption The biggest obstacle to commercial production of LLM is the problem of illusion.

Mitigating the phenomenon of hallucinations and developing metrics to measure them is a booming research topic, with many startups focused on solving this problem.

There are currently some temporary methods to reduce hallucinations, such as adding more context, thought chains, self-consistency to prompts, or requiring the output of the model to remain concise.

The following are related speeches that you can refer to

·Survey of Hallucination in Natural Language Generation (Ji et al., 2022)·How Language Model Hallucinations Can Snowball (Zhang et al., 2023)·A Multitask, Multilingual, Multimodal uation of ChatGPT on Reasoning, Hallucination, and Interactivity (Bang et al., 2023)·Contrastive Learning Reduces Hallucination in Conversations (Sun et al., 2022)·Self-Consistency Improves Chain of Thought Reasoning in Language Models (Wang et al., 2022)·SelfCheckGPT: Zero-Resource Black-Box Hallucination Detection for Generative Large Language Models (Manakul et al., 2023)

02 Optimize context length and context construction

The vast majority of problems faced by AI require context.

For example, if we ask ChatGPT: "Which Vietnamese restaurant is the best?", the required context might be "where" because the best restaurant in Vietnam may be different from the best Vietnamese restaurant in the United States.

According to the interesting paper "SituatedQA" (Zhang & Choi, 2021), a considerable proportion of information-seeking questions have context-dependent answers. For example, about 16.5% of the questions in the NQ-Open dataset are of this type. .

I personally think that for enterprise application scenarios, this ratio may be even higher. Suppose a company builds a chatbot for customers. If the robot is to be able to answer any customer question about any product, the required context may be the customer's history or information about the product.

Because the model "learns" from the context provided to it, this process is also known as contextual learning.

For retrieval enhanced generation (RAG, which is also the main method in the LLM industry application direction), context length is particularly important.

RAG can be simply divided into two stages:

Phase 1: Chunking (also called indexing)

Collect all documents to be used by LLM, split these documents into chunks that can be fed into LLM to generate embeddings, and store these embeddings in a vector database.

Second stage: query

When a user sends a query, such as “Will my insurance policy cover this drug

Figure: Screenshot from Jerry Liu’s speech on LlamaIndex (2023)

The longer the context length, the more blocks we can insert into the context. But will the more information a model has access to, the better its responses will be?

This isn't always the case. How much context a model can use and how efficiently the model will be used are two different questions. Just as important as increasing the model context length is more efficient learning of the context, which is also called "hint engineering".

A recent widely circulated paper shows that models perform much better at understanding information from the beginning and end of the index than from the middle: Lost in the Middle: How Language Models Use Long Contexts (Liu et al., 2023).

03Integrate other data modes

In my opinion, multimodality is so powerful yet often underestimated.

First of all, many real-life application scenarios require processing large amounts of multi-modal data, such as healthcare, robotics, e-commerce, retail, games, entertainment, etc. Medical predictions require the use of both text (such as doctor's notes, patient questionnaires) and images (such as CT, X-ray, MRI scans); product data often includes images, videos, descriptions, and even tabular data (such as production date, weight, color).

Second, multimodality promises to bring huge improvements in model performance. Wouldn’t a model that could understand both text and images perform better than a model that could only understand text? Text-based models require large amounts of text data, and now we are really worried about running out of internet data for training text-based models. Once the text is exhausted, we need to leverage other data modalities.

One application direction that I'm particularly excited about recently is that multimodal technology can help visually impaired people browse the Internet and navigate the real world.

The following are several outstanding multimodal research developments:· [CLIP] Learning Transferable Visual Models From Natural Language Supervision (OpenAI, 2021)·Flamingo: a Visual Language Model for Few-Shot Learning (DeepMind, 2022)·BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models (Salesforce, 2023)·KOSMOS-1: Language Is Not All You Need: Aligning Perception with Language Models (Microsoft, 2023)·PaLM-E: An embodied multimodal language model (Google, 2023)·LLaVA: Visual Instruction Tuning (Liu et al., 2023)·NeVA: NeMo Vision and Language Assistant (NVIDIA, 2023)

04Improving the speed and reducing costs of LLMs

When GPT-3.5 was first launched in late November 2022, many people expressed concerns about the delays and costs of using the model in production.

Now, the delay/cost analysis caused by the use of GPT-3.5 has taken a new turn. Within half a year, the entire modeling community found a new way to create a model that was almost close to GPT-3.5 in terms of performance, but with less than 2% of the memory footprint.

One of my points from this is: if you create something good enough, someone else will find a way to make it fast and cost-effective.

The following is based on data reported in the Guanaco paper, which compares the performance of Guanaco 7B with ChatGPT GPT-3.5 and GPT-4.

It is important to note that, overall, the performance of these models is far from perfect. For LLM, it is still very difficult to significantly improve performance.

I remember four years ago, when I started writing the notes for the "Model Compression" section of the book "Designing Machine Learning Systems", there were four main model optimization/compression techniques in the industry:

  1. Quantification: by far the most common model optimization method. Quantization reduces the size of the model by using fewer bits to represent the parameters of the model. For example, instead of using 32 bits to represent floating point numbers, only 16 bits or even 4 bits are used.

  2. Knowledge distillation: that is, training a small model (student model), which can imitate a larger model or model set (teacher model).

  3. Low-rank decomposition: Its key idea is to use low-dimensional tensors to replace high-dimensional tensors to reduce the number of parameters. For example, a 3x3 tensor can be decomposed into the product of a 3x1 tensor and a 1x3 tensor, so that instead of 9 parameters, there are only 6 parameters.

  4. Pruning: refers to reducing the size of the model by removing weights or connections in the model that contribute less to the overall performance.

These four techniques are still popular today. Alpaca is trained through knowledge distillation, while QLoRA uses a combination of low-rank decomposition and quantization.

05Design new model architecture

Since AlexNet in 2012, we have seen many architectures come and go, including LSTM, seq2seq, etc.

Compared with these architectures, Transformer, which was launched in 2017, is extremely stable, although it is unclear how long this architecture will be popular.

It is not easy to develop a new architecture that can outperform Transformer. In the past 6 years, Transformer has undergone a lot of optimization. On suitable hardware, the scale and effect of this model can achieve amazing results (PS: Transformer was first designed by Google to run quickly on TPU , and was later optimized on the GPU).

In 2021, the research "Efficiently Modeling Long Sequences with Structured State Spaces" (Gu et al., 2021) by Chris Ré's laboratory triggered a lot of discussions in the industry. I'm not sure what happened next. But Chris Ré Labs is still actively developing new architectures, and they recently launched an architecture called Monarch Mixer in partnership with startup Together.

Their main idea is that for the existing Transformer architecture, the complexity of attention is proportional to the square of the sequence length, and the complexity of MLP is proportional to the square of the model dimension. Architectures with sub-quadratic complexity will be more efficient.

I'm sure many other labs are exploring this idea, although I'm not aware of any studies that have publicly tried it. If you know the progress, please contact me!

06Developing GPU Alternatives

Since the advent of AlexNet in 2012, GPU has been the main hardware for deep learning.

In fact, one of the generally recognized reasons for AlexNet’s popularity is that it was the first paper to successfully use GPUs to train neural networks. Before GPUs, if you wanted to train a model of the size of AlexNet, you would need thousands of CPUs, just like the server Google released a few months before AlexNet.

Compared with thousands of CPUs, a few GPUs are more accessible to PhD students and researchers, triggering a boom in deep learning research.

Over the past decade, many companies, both large and startups, have attempted to create new hardware for artificial intelligence. The most noteworthy attempts include Google's TPU, Graphcore's IPU, and Cerebras. SambaNova has also raised more than $1 billion to develop new AI chips, but appears to have pivoted to becoming a generative AI platform.

During this period, quantum computing also aroused a lot of expectations, among which the main players include:

·IBM’s quantum processor

·Google’s quantum computer. A major milestone in quantum error reduction was reported in Nature earlier this year. Its quantum virtual machine is publicly accessible through Google Colab.

·Research laboratories in universities, such as MIT Quantum Engineering Center, Max Planck Institute for Quantum Optics, Chicago Quantum Exchange Center, etc.

Another equally exciting direction is photonic chips. This is the direction I know the least about. If there are any mistakes, please correct me.

Existing chips use electricity to transmit data, which consumes a lot of energy and creates latency. Photonic chips use photons to transmit data, harnessing the speed of light for faster, more efficient computing. Various startups in this space have raised hundreds of millions of dollars, including Lightmatter ($270 million), Ayar Labs ($220 million), Lightelligence ($200 million+), and Luminous Computing ($115 million).

The following is the progress timeline of the three main methods of photon matrix calculation, excerpted from Photonic matrix multiplication lights up photonic accelerator and beyond (Zhou et al., Nature 2022). The three different methods are Planar Light Conversion (PLC), Mach-Zehnder Interferometer (MZI) and Wavelength Division Multiplexing (WDM).

07Improving agent availability

Agents can be thought of as LLMs that can take actions, such as browsing the Internet, sending emails, etc. Compared with other research directions in this article, this may be the youngest direction.

There is great interest in agents due to their novelty and great potential. Auto-GPT is now the 25th most popular library by number of stars on GitHub. GPT-Engineering is also another popular library.

Despite this, there are still doubts about whether LLMs are reliable enough, perform well enough, and have certain operational capabilities.

Now there is an interesting application direction, which is to use agents for social research. A Stanford experiment showed that a small group of generative agents produced emergent social behavior: starting with just one user-specified idea, that one agent wanted to host a Valentine's Day party, a number of other agents spread it autonomously over the next two days. Invitations to parties, making new friends, inviting each other to parties...(Generative Agents: Interactive Simulacra of Human Behavior, Park et al., 2023).

Perhaps the most noteworthy startup in this space is Adept, founded by two Transformer co-authors (although both have since left) and a former OpenAI VP, and which has raised nearly $500 million to date Dollar. Last year, they showed how their agent could browse the internet and add new accounts on Salesforce. I look forward to seeing their new demo 🙂 .

08 Improving the ability to learn from human preferences

RLHF (Reinforcement Learning from Human Preference) is cool, but a bit tedious.

I'm not surprised that people will find better ways to train LLMs. There are many open questions regarding RLHF, such as:

·How to represent human preferences mathematically?

Currently, human preferences are determined through comparison: a human annotator determines whether answer A is better than answer B. However, it does not take into account the specific extent to which answer A is better or worse than answer B.

·What are human preferences?

Anthropic measures the quality of model responses along three dimensions: helpful, honest, and harmless. Reference paper: Constitutional AI: Harmlessness from AI Feedback (Bai et al., 2022).

DeepMind tries to generate answers that will best please the most people. Reference paper: Fine-tuning language models to find agreement among humans with diverse preferences, (Bakker et al., 2022).

Also, do we want an AI that can take a stand, or a generic AI that avoids talking about any potentially controversial topic?

·Whose preferences are “human” preferences, taking into account differences in culture, religion, political leanings, etc.?

There are many challenges in obtaining training data that is sufficiently representative of all potential users.

For example, OpenAI's InstructGPT data has no annotators over 65 years old. The taggers are mainly Filipinos and Bangladeshis. Reference paper: InstructGPT: Training language models to follow instructions with human feedback (Ouyang et al., 2022).

Although the original intentions of AI community-led efforts in recent years are admirable, data bias still exists. For example, in the OpenAssistant dataset, 201 of 222 respondents (90.5%) self-reported as male. Jeremy Howard posted a series of tweets about the issue on Twitter.

09Improve the efficiency of the chat interface

Since the introduction of ChatGPT, there has been an ongoing discussion about whether chat is suitable for a wide range of tasks. for example:

·Natural language is the lazy user interface (Austin Z. Henley, 2023)

·Why Chatbots Are Not the Future (Amelia Wattenberger, 2023)

·What Types of Questions Require Conversation to Answer? A Case Study of AskReddit Questions (Huang et al., 2023)

·AI chat interfaces could become the primary user interface to read documentation (Tom Johnson, 2023)

·Interacting with LLMs with Minimal Chat (Eugene Yan, 2023)

However, this is not a new discussion. In many countries, especially in Asia, chat has been used as the interface for super apps for about a decade. Dan Grover discussed this phenomenon in 2014.

This type of discussion became hot again in 2016, with many people taking the view that existing application types are obsolete and that chatbots are the future. For example, the following studies:

·On chat as interface (Alistair Croll, 2016)

·Is the Chatbot Trend One Big Misunderstanding? (Will Knight, 2016)

·Bots won’t replace apps. Better apps will replace apps (Dan Grover, 2016)

Personally, I like the chat interface for the following reasons:

The chat interface is one that everyone (even people with no previous experience with computers or the Internet) can quickly learn to use.

When I was volunteering in a low-income neighborhood in Kenya in the early 2010s, I was surprised to see how comfortable everyone there was with banking via text message on their phone. Even if no one in that community has a computer.

The chat interface is generally easy to access. We can also use speech instead of text if our hands are busy with other things.

The chat interface is also a very powerful interface. It will respond to any request made by the user, even if some of the responses are not very good.

However, I think there are some areas where the chat interface could be improved:

·Multiple messages in one round

Currently, we pretty much assume there is only one message at a time. But when my friends and I text, it often takes multiple messages to complete a chat because I need to insert different data (e.g. images, locations, links), I forgot something from a previous message, or I just don't want to fit everything into one big paragraph.

·Multimodal input

In the field of multimodal applications, most efforts are spent on building better models and less on building better interfaces. Take NVIDIA’s NeVA chatbot as an example. I'm no user experience expert, but I think there might be room for improvement here.

PS Sorry, NeVA team, for naming you. Still, your work is awesome!

Figure: NVIDIA’s NeVA interface

·Integrate generative AI into workflows

Linus Lee covers this very well in his talk "Generative AI interface beyond chats". For example, if you want to ask a question about a chart column you're working on, you should be able to just point to that column and ask.

·Edit and delete messages

How does editing or removing user input change the flow of the conversation with the chatbot?

10 Building LLMs for non-English languages

We know that current English-led LLMs perform poorly in many other languages, whether in terms of performance, latency or speed.

Here are relevant studies you can refer to:

·ChatGPT Beyond English: Towards a Comprehensive uation of Large Language Models in Multilingual Learning (Lai et al., 2023)

·All languages are NOT created (tokenized) equal (Yennie Jun, 2023)

Some readers have told me that they don't think I should pursue this direction for two reasons.

This is more of a "logistical" question than a research question. We already know how to do it. Someone just needs to put in the money and effort.

This is not entirely correct. Most languages are considered low-resource languages, as they have much less high-quality data than English or Chinese, for example, and may require different techniques to train large language models.

Here are relevant studies you can refer to:

·Low-resource Languages: A Review of Past Work and Future Challenges (Magueresse et al., 2020)

·JW300: A Wide-Coverage Parallel Corpus for Low-Resource Languages (Agić et al., 2019)

Those who are more pessimistic believe that in the future, many languages will die out and the Internet will be two worlds composed of two languages: English and Chinese. This way of thinking is not new. Does anyone remember Esperanto?

The impact of AI tools, such as machine translation and chatbots, on language learning remains unclear. Will they help people learn new languages faster, or will they eliminate the need to learn new languages altogether?

in conclusion

Of the 10 challenges mentioned above, some are indeed more difficult than others.

For example, I think item 10, Building LLMs for non-English languages, more directly points to adequate time and resources.

Item 1, reducing hallucinations, will be more difficult because hallucinations are just LLMs doing their probabilistic task.

Item 4, making LLMs faster and cheaper, will never reach a fully solved state. A lot of progress has been made in this area and there is more to come, but we will never stop improving.

Items 5 and 6, new architecture and new hardware, are very challenging and inevitable. Due to the symbiotic relationship between architecture and hardware, new architectures need to be optimized for common hardware, and hardware needs to support common architectures. They may be settled by the same company.

Some of these problems can be solved with more than just technical knowledge. For example, Item 8, Improving Learning from Human Preferences, may be more of a strategy issue than a technical issue.

Item 9, improving the efficiency of the chat interface, is more of a user experience issue. We need more people with non-technical backgrounds working together to solve these problems.

Disclaimer: The information on this page may come from third-party sources and is for reference only. It does not represent the views or opinions of Gate and does not constitute any financial, investment, or legal advice. Virtual asset trading involves high risk. Please do not rely solely on the information on this page when making decisions. For details, see the Disclaimer.
Comment
0/400
No comments