How a Browser Work
What a browser actually is (beyond “it opens websites”)
A browser is a client-side application that communicates with web servers using standardized protocols, retrieves resources, interprets them, executes code, enforces security rules, and renders visual output for human interaction.
A web browser is a secure client-side software platform that communicates with servers using internet protocols, retrieves and interprets web resources, executes scripts, enforces security policies, and renders interactive content for users.
1. Browser as a Client in the Client–Server Model
2. Browser as a Network Stack User
3. Browser as a Resource Fetcher (Not a Page Loader)
4. Browser as a Parser and Interpreter
5. Browser as a Rendering Engine
6. Browser as a JavaScript Runtime Environment
7. Browser as a Security Enforcer
8. Browser as a Process Manager
Main parts of a browser (high-level overview)
A modern web browser is made up of several major components, each with a specific responsibility. Together, they turn raw internet data into an interactive webpage.
A browser consists of the user interface, browser engine, rendering engine, JavaScript engine, networking layer, UI backend, data storage, and security system, all working together to fetch, process, and display web content securely.
1. User Interface (UI)
2. Browser Engine
3. Rendering Engine
4. JavaScript Engine
5. Networking Layer
6. UI Backend / Graphics Layer
7. Data Storage / Persistence
8. Security & Sandbox System
User Interface: address bar, tabs, buttons
The User Interface is the visible and interactive part of a web browser. It is the layer through which users control the browser and navigate the web. Although it looks simple, it plays a crucial coordinating role.
1. Address Bar (URL Bar / Omnibox)
The address bar is the primary input point of the browser.
Functions:
Accepts website addresses (URLs)
Accepts search queries
Shows the current website’s address
Displays security information (HTTPS lock icon)
Warns about unsafe sites
2. Tabs
Tabs allow users to open multiple webpages within a single browser window.
Functions:
Manage multiple sessions simultaneously
Enable fast switching between pages
Isolate pages for stability and security
3. Navigation Buttons (Back, Forward, Refresh, Home)
These buttons control page navigation.
Back / Forward
Move through browsing history
Use cached pages when possible
Do not always trigger full reloads
The User Interface does not render web pages.
It only collects user actions and passes instructions to the browser engine. The user interface of a browser includes the address bar, tabs, and navigation buttons, allowing users to enter URLs, manage multiple pages, and control navigation while coordinating with the browser engine.
Browser Engine vs Rendering Engine (simple distinction)
Browser engine controls the process; rendering engine controls the appearance. The browser engine manages user actions and controls page loading, while the rendering engine interprets HTML and CSS to display the webpage on the screen.
Browser Engine
What it is:
The controller / manager of the browser.
What it does:
Takes input from the User Interface (URL bar, buttons, tabs)
Decides what page to load
Coordinates between:
Rendering engine
JavaScript engine
Networking layer
Handles navigation (back, forward, reload)
Think of it as:
🧠 The boss that gives instructions.
Rendering Engine
What it is:
The worker that turns code into visuals.
What it does:
Parses HTML
Parses CSS
Builds DOM and CSSOM
Calculates layout
Paints pixels on the screen
Think of it as:
🎨 The painter / builder that draws the webpage.
Difference Between Browser Engine & Rendering Engine
| Browser Engine | Rendering Engine |
| Manages browser actions | Draws the webpage |
| Handles navigation | Handles layout & paint |
| Talks to UI & network | Talks to graphics layer |
| Decides what to load | Decides how it looks |
Networking: how a browser fetches HTML, CSS, JS
Fetching the HTML
The browser sends an HTTP request:
GET / HTTP/1.1
Host: example.com
The server responds with:
HTML document
HTTP headers (status code, content type, caching rules)
➡️ HTML is always fetched first because it describes the structure of the page.
Fetching CSS Files
For every CSS file:
Browser sends a separate HTTP request
CSS is downloaded and parsed
CSS blocks rendering until styles are known
➡️ CSS is critical because layout depends on it.
Fetching JavaScript Files
For JavaScript:
Each
<script>causes a requestBy default, JS blocks HTML parsing
Browser may pause rendering to execute JS
HTML parsing and DOM creation
HTML Parsing and DOM Creation
Once the browser has started fetching the HTML, the next critical job is to parse that HTML and convert it into a structure the browser can work with. That structure is called the DOM (Document Object Model).
What HTML Parsing Means
HTML parsing is the process where the browser:
Reads raw HTML text
Understands tags and elements
Converts them into objects in memory
Parsing Happens Incrementally (Streaming)
Important point:
The browser does not wait for the full HTML file
It parses HTML as it arrives from the network
Tokenization (Breaking HTML into Tokens)
The parser first tokenizes the HTML.
Tokens created:
Start tag:
<p>Text:
HelloEnd tag:
</p>
The HTML DOM (Document Object Model) is a structured representation of a web page that allows developers to access, modify, and control its content and structure using JavaScript. It powers most dynamic website interactions, enabling features like real-time updates, form validation, and interactive user interfaces.
The HTML Document Object Model (DOM) is a tree structure, where each HTML tag becomes a node in the hierarchy.
At the top, the <html> tag is the root element, containing both <head> and<body> as child elements.
These in turn have their own children, such as <title>, <div>, and nested elements like <h1>, <p>,<ul>, and<li>.
Elements that contain other elements are labeled as container elements, while elements that do not are simply child elements.
This hierarchy allows developers to navigate and manipulate web page content using JavaScript by traversing from parent to child and vice versa.
CSS parsing and CSSOM creation
CSS parsing is the vital initial step in the process of browser rendering and applying CSS styles to HTML elements. It involves analyzing and breaking down CSS code into a structured format that can be understood by the browser rendering engine.
CSSOM creation is the process by which the browser converts CSS rules into a structured object model that represents all the styles applied to a webpage.
How CSSOM Is Created
Step 1: CSS Is Downloaded
CSS can come from:
External files (
<link rel="stylesheet">)<style>blocksInline styles
Step 2: Tokenization
The CSS parser breaks CSS into tokens:
Selectors
Properties
Values
Step 3: Rule Parsing
Tokens are converted into CSS rules:
Selector (
p)Declarations (
color,font-size)
Step 4: CSSOM Tree Construction
All CSS rules are organized into a tree-like structure.
The CSSOM:
Represents the cascade
Stores inheritance relationships
Handles specificity and order
How DOM and CSSOM come together
After the browser has parsed the HTML and parsed the CSS, it still cannot display anything.
This is because structure alone (DOM) and styles alone (CSSOM) are not enough.
The browser must combine both to decide what to draw and how to draw it.Separate Creation of DOM and CSSOM
DOM (Document Object Model)
Created from HTML
Represents the structure and content of the webpage
Stored as a tree of nodes (elements, text, etc.)
CSSOM (CSS Object Model)
Created from CSS
Represents all styling rules
Includes:
External stylesheets
Internal
<style>rulesInline styles
Default browser styles
After parsing HTML and CSS, the browser combines the DOM and CSSOM during style calculation to apply CSS rules to DOM elements. This results in the render tree, which includes only visible elements with computed styles. The render tree is then used for layout calculation and painting, allowing the browser to display the webpage correctly.
Layout (reflow), painting, and display
After the browser has combined the DOM and CSSOM into the Render Tree, it still does not immediately show anything.
The browser must now decide where each element goes, how it looks visually, and finally display it on the screen.
This happens in three major stages:
Layout (Reflow)
Painting
Display (Compositing)
1. Layout (Reflow)
What Layout Means:
Layout, also called reflow, is the process where the browser calculates the exact size and position of every element in the render tree.
2. Painting
What Painting Means
After layout is complete, the browser draws the visual appearance of elements.
The browser paints:
Text
Colors
Backgrounds
Borders
Images
Shadows
Gradients
3. Display (Compositing)
What Display / Compositing Means
After painting:
The browser combines painted layers
Sends them to the GPU
Displays the final image on the screen
Difference Between Layout, Paint, and Display
| Stage | What it decides | Cost |
| Layout (Reflow) | Size & position | High |
| Painting | Visual appearance | Medium |
| Display (Composite) | Final screen output | Low |