What is a large language model?
How an LLM works at a high level, what it is good at and where it fails.
A large language model (LLM) is a neural network trained on huge amounts of text to predict what comes next. It reads text as a sequence of tokens and generates a reply one token at a time. Chatbots, coding assistants and summarizers are built on top of models like this.
What that means in practice
- It is good at language tasks: summarizing, rewriting, translating, drafting, classifying and explaining code.
- It generates plausible text. It does not look facts up unless you give it tools or documents, so it can state wrong things confidently (often called hallucinations).
- It has no memory between requests. Whatever it should know must be sent again in each request.
Chat apps vs APIs
A chat app is a product with a fixed price per month. An API lets your own software send text to a model and receive the answer, billed by usage in tokens. That is what this course is about.
A request, step by step
- Your program sends the input: instructions, any documents and the conversation so far.
- The provider runs the model.
- It returns the output text.
- You are billed for the input tokens and the output tokens.
NoteBecause every request is separate, a long conversation gets more expensive each turn: the whole history is sent again as input.
Test yourself
Answer all the questions, then check them. Finish with every answer right to mark the lesson as done.