Skip to main content

Command Palette

Search for a command to run...

How a Browser Work

Published
•8 min read•View as Markdown

What a browser actually is (beyond “it opens websites”)

A browser is a client-side application that communicates with web servers using standardized protocols, retrieves resources, interprets them, executes code, enforces security rules, and renders visual output for human interaction.

A web browser is a secure client-side software platform that communicates with servers using internet protocols, retrieves and interprets web resources, executes scripts, enforces security policies, and renders interactive content for users.

1. Browser as a Client in the Client–Server Model

2. Browser as a Network Stack User

3. Browser as a Resource Fetcher (Not a Page Loader)

4. Browser as a Parser and Interpreter

5. Browser as a Rendering Engine

6. Browser as a JavaScript Runtime Environment

7. Browser as a Security Enforcer

8. Browser as a Process Manager

Main parts of a browser (high-level overview)

A modern web browser is made up of several major components, each with a specific responsibility. Together, they turn raw internet data into an interactive webpage.

A browser consists of the user interface, browser engine, rendering engine, JavaScript engine, networking layer, UI backend, data storage, and security system, all working together to fetch, process, and display web content securely.

1. User Interface (UI)

2. Browser Engine

3. Rendering Engine

4. JavaScript Engine

5. Networking Layer

6. UI Backend / Graphics Layer

7. Data Storage / Persistence

8. Security & Sandbox System

User Interface: address bar, tabs, buttons

The User Interface is the visible and interactive part of a web browser. It is the layer through which users control the browser and navigate the web. Although it looks simple, it plays a crucial coordinating role.

1. Address Bar (URL Bar / Omnibox)

The address bar is the primary input point of the browser.

Functions:

  • Accepts website addresses (URLs)

  • Accepts search queries

  • Shows the current website’s address

  • Displays security information (HTTPS lock icon)

  • Warns about unsafe sites

2. Tabs

Tabs allow users to open multiple webpages within a single browser window.

Functions:

  • Manage multiple sessions simultaneously

  • Enable fast switching between pages

  • Isolate pages for stability and security

3. Navigation Buttons (Back, Forward, Refresh, Home)

These buttons control page navigation.

Back / Forward

  • Move through browsing history

  • Use cached pages when possible

  • Do not always trigger full reloads

The User Interface does not render web pages.
It only collects user actions and passes instructions to the browser engine. The user interface of a browser includes the address bar, tabs, and navigation buttons, allowing users to enter URLs, manage multiple pages, and control navigation while coordinating with the browser engine.

Browser Engine vs Rendering Engine (simple distinction)

Browser engine controls the process; rendering engine controls the appearance. The browser engine manages user actions and controls page loading, while the rendering engine interprets HTML and CSS to display the webpage on the screen.

Browser Engine

What it is:
The controller / manager of the browser.

What it does:

  • Takes input from the User Interface (URL bar, buttons, tabs)

  • Decides what page to load

  • Coordinates between:

    • Rendering engine

    • JavaScript engine

    • Networking layer

  • Handles navigation (back, forward, reload)

Think of it as:
🧠 The boss that gives instructions.

Rendering Engine

What it is:
The worker that turns code into visuals.

What it does:

  • Parses HTML

  • Parses CSS

  • Builds DOM and CSSOM

  • Calculates layout

  • Paints pixels on the screen

Think of it as:
🎨 The painter / builder that draws the webpage.

Difference Between Browser Engine & Rendering Engine

Browser EngineRendering Engine
Manages browser actionsDraws the webpage
Handles navigationHandles layout & paint
Talks to UI & networkTalks to graphics layer
Decides what to loadDecides how it looks

Networking: how a browser fetches HTML, CSS, JS

Fetching the HTML

The browser sends an HTTP request:

GET / HTTP/1.1

Host: example.com

The server responds with:

  • HTML document

  • HTTP headers (status code, content type, caching rules)

➡️ HTML is always fetched first because it describes the structure of the page.

Fetching CSS Files

For every CSS file:

  • Browser sends a separate HTTP request

  • CSS is downloaded and parsed

  • CSS blocks rendering until styles are known

➡️ CSS is critical because layout depends on it.

Fetching JavaScript Files

For JavaScript:

  • Each <script> causes a request

  • By default, JS blocks HTML parsing

  • Browser may pause rendering to execute JS

    HTML parsing and DOM creation

    HTML Parsing and DOM Creation

    Once the browser has started fetching the HTML, the next critical job is to parse that HTML and convert it into a structure the browser can work with. That structure is called the DOM (Document Object Model).

    steps in HTML parsing process

    What HTML Parsing Means

    HTML parsing is the process where the browser:

    • Reads raw HTML text

    • Understands tags and elements

    • Converts them into objects in memory

Parsing Happens Incrementally (Streaming)

Important point:

  • The browser does not wait for the full HTML file

  • It parses HTML as it arrives from the network

Tokenization (Breaking HTML into Tokens)

The parser first tokenizes the HTML.

Tokens created:

  • Start tag: <p>

  • Text: Hello

  • End tag: </p>

The HTML DOM (Document Object Model) is a structured representation of a web page that allows developers to access, modify, and control its content and structure using JavaScript. It powers most dynamic website interactions, enabling features like real-time updates, form validation, and interactive user interfaces.

  • The HTML Document Object Model (DOM) is a tree structure, where each HTML tag becomes a node in the hierarchy.

  • At the top, the <html> tag is the root element, containing both <head> and<body> as child elements.

  • These in turn have their own children, such as <title>, <div>, and nested elements like <h1>, <p>,<ul>, and<li>.

  • Elements that contain other elements are labeled as container elements, while elements that do not are simply child elements.

  • This hierarchy allows developers to navigate and manipulate web page content using JavaScript by traversing from parent to child and vice versa.

    CSS parsing and CSSOM creation

    CSS parsing is the vital initial step in the process of browser rendering and applying CSS styles to HTML elements. It involves analyzing and breaking down CSS code into a structured format that can be understood by the browser rendering engine.

css parsing steps diagram

CSSOM creation is the process by which the browser converts CSS rules into a structured object model that represents all the styles applied to a webpage.

How CSSOM Is Created

Step 1: CSS Is Downloaded

CSS can come from:

  • External files (<link rel="stylesheet">)

  • <style> blocks

  • Inline styles

Step 2: Tokenization

The CSS parser breaks CSS into tokens:

  • Selectors

  • Properties

  • Values

Step 3: Rule Parsing

Tokens are converted into CSS rules:

  • Selector (p)

  • Declarations (color, font-size)

Step 4: CSSOM Tree Construction

All CSS rules are organized into a tree-like structure.

The CSSOM:

  • Represents the cascade

  • Stores inheritance relationships

  • Handles specificity and order

    How DOM and CSSOM come together

    After the browser has parsed the HTML and parsed the CSS, it still cannot display anything.
    This is because structure alone (DOM) and styles alone (CSSOM) are not enough.
    The browser must combine both to decide what to draw and how to draw it.

    Separate Creation of DOM and CSSOM

    DOM (Document Object Model)

    • Created from HTML

    • Represents the structure and content of the webpage

    • Stored as a tree of nodes (elements, text, etc.)

CSSOM (CSS Object Model)

  • Created from CSS

  • Represents all styling rules

  • Includes:

    • External stylesheets

    • Internal <style> rules

    • Inline styles

    • Default browser styles

After parsing HTML and CSS, the browser combines the DOM and CSSOM during style calculation to apply CSS rules to DOM elements. This results in the render tree, which includes only visible elements with computed styles. The render tree is then used for layout calculation and painting, allowing the browser to display the webpage correctly.

Layout (reflow), painting, and display

After the browser has combined the DOM and CSSOM into the Render Tree, it still does not immediately show anything.
The browser must now decide where each element goes, how it looks visually, and finally display it on the screen.

This happens in three major stages:

  1. Layout (Reflow)

  2. Painting

  3. Display (Compositing)

1. Layout (Reflow)

What Layout Means:

Layout, also called reflow, is the process where the browser calculates the exact size and position of every element in the render tree.

2. Painting

What Painting Means

After layout is complete, the browser draws the visual appearance of elements.

The browser paints:

  • Text

  • Colors

  • Backgrounds

  • Borders

  • Images

  • Shadows

  • Gradients

3. Display (Compositing)

What Display / Compositing Means

After painting:

  • The browser combines painted layers

  • Sends them to the GPU

  • Displays the final image on the screen

Difference Between Layout, Paint, and Display

StageWhat it decidesCost
Layout (Reflow)Size & positionHigh
PaintingVisual appearanceMedium
Display (Composite)Final screen outputLow