Posts

Why Are Interrupts Better Than Polling?

Image
In embedded systems, peripherals constantly generate events. A UART receives a byte. A timer expires. A GPIO pin changes state. An ADC conversion completes. The processor needs to respond to these events. There are two common approaches: Polling Interrupts At first, polling looks simple. But interrupts usually provide a more efficient way for the CPU to respond to asynchronous events. What Is Polling? In polling, the CPU repeatedly checks whether an event has occurred. For example: while (1) { if (UART_RX_READY) { read_uart_data(); } } The processor keeps asking the UART: "Did data arrive?" If no data has arrived, the CPU checks again. And again. And again. Imagine UART data arrives only once every 100 ms. During those 100 ms, the CPU may check the same status flag thousands of times even though nothing has happened. That CPU time could have been used for something else. What Changes With Interrupts? With interrupt...

Why Does the Stack Usually Grow Downward?

Image
When debugging nested function calls, you may notice the Stack Pointer moving toward lower memory addresses as the call depth increases. For example: SP = 0x20001000 PUSH {R4} SP = 0x20000FFC But why does this happen? Is a downward-growing stack inherently better, or is it simply an architectural convention? 1. The Stack Does Not Have to Grow Downward A processor can support either: Descending stack: SP decreases as data is pushed. or: Ascending stack: SP increases as data is pushed. Both approaches can work efficiently. So the real question is: Why did descending stacks become so common? 2. It Fits Traditional Memory Layouts Well A simplified memory layout often looks like this: High Address +------------------+ | Stack | | ↓ | | | | Free RAM | | | | ↑ | | Heap | +------------------+ Low Address The stack can begin near the upper end of available memory and grow toward lower addr...

Why Do Microcontrollers Start with an Internal Oscillator?

Why Does a Microcontroller Start With an Internal Oscillator? When a microcontroller powers on, it needs one thing before it can execute even its first instruction: A clock. Most microcontrollers initially run from an internal oscillator. But if an external crystal is more accurate, why doesn't the microcontroller simply start with it? The answer comes down to startup speed, reliability, and initialization sequence. 1️⃣ Fast Startup Internal RC oscillators can become usable very quickly. External crystals need additional time for their oscillations to build and stabilize. Instead of waiting, the MCU can start executing firmware using its internal clock. Power ON → Internal Oscillator → CPU Starts Executing Meanwhile, the external clock source can be initialized and allowed to stabilize. 2️⃣ Guaranteed Availability The internal oscillator is built into the microcontroller. An external crystal depends on components outside the MCU, such as: • Crystal or resonator • Load capacitors • ...

๐Ÿค” Why Is RAM Access Slower Than CPU Registers in ARM Microcontrollers?

CPU registers are located inside the processor core . RAM is located outside the CPU core . So accessing RAM requires: address generation bus access memory read/write cycles But registers can be accessed almost instantly by the CPU. That is why CPUs first load data into registers before processing. ๐Ÿ‘‰ Faster register access = better execution speed. This is also why: repeated memory access slows firmware efficient register usage improves performance compilers try to keep frequently used variables in registers Understanding this helps in: firmware optimization driver development assembly understanding performance tuning Small low-level concept. Huge embedded impact. #TheWhySeries #EmbeddedSystems #Firmware #ARM #Microcontrollers #EmbeddedC

๐Ÿš€ Why Do MCUs Have Vector Tables?

When an interrupt or reset occurs, the CPU must quickly know which function to execute. Instead of searching through code, MCUs use a Vector Table — a fixed memory table that stores the addresses of interrupt handlers. Example (ARM Cortex-M): C void Reset_Handler(void); void UART_IRQHandler(void); void SysTick_Handler(void); The vector table stores pointers to these handlers: 0x00000000 → Initial Stack Pointer   0x00000004 → Reset_Handler   0x00000008 → NMI_Handler   0x0000000C → HardFault_Handler ... When an interrupt occurs, the CPU simply reads the handler address from the vector table and jumps to it. ๐Ÿ“Œ Why this design? • Instant interrupt response • Simple hardware implementation • Deterministic interrupt latency ๐Ÿ’ก Key Insight The vector table acts like a hardware lookup table that maps interrupts to their handlers. #EmbeddedSystems #Firmware #Microcontrollers #ComputerArchitecture #WhySeries

๐Ÿš€ Why Do CPUs Perform Operations on Registers Instead of Memory?

In most processors, the ALU (Arithmetic Logic Unit) operates only on CPU registers, not directly on memory. Example: Asm LDR R1, [R0] ADD R2, R1, R3 STR R2, [R0] The CPU first loads data from memory into registers, performs the operation, and then stores the result back. But why? ๐Ÿ“Œ Speed Registers are inside the CPU core, so ALU operations can often complete in a single clock cycle. ๐Ÿ“Œ Memory is slower Accessing RAM involves the address bus, memory controller, and data bus, which takes many cycles. ๐Ÿ“Œ Simpler CPU design Keeping ALU operations on registers allows faster pipelines and predictable instruction timing. ๐Ÿ’ก Key Insight Efficient firmware minimizes memory access and performs as many operations as possible in registers. That’s why optimized drivers often read a register once, modify it in CPU registers, and write it back. #EmbeddedSystems #Firmware #ComputerArchitecture #Microcontrollers #EmbeddedLearning #WhySeries

❓ Why Do Microcontrollers Use Memory-Mapped I/O?

In most microcontrollers, peripherals like GPIO, UART, SPI, and Timers are accessed as if they were normal memory locations. Example: C GPIO->OUT |= LED1; Behind the scenes, the CPU is simply reading or writing a specific memory address assigned to that register. Example memory map: 0x00000000 → Flash   0x20000000 → SRAM   0x40000000 → Peripherals So when firmware accesses: 0x40020014 the CPU is actually talking to a GPIO register, not RAM. Why this design? ✔ Simplifies CPU design — same instructions for memory and peripherals ✔ Allows standard load/store instructions to control hardware ✔ Makes firmware development easier ✔ Enables compilers to generate simple and efficient code ๐Ÿ’ก Key insight Peripherals are not accessed with special instructions. They are simply memory addresses mapped to hardware registers. That’s why embedded firmware often looks like: Memory → Register → Processing → Memory #EmbeddedSystems #Firmware #Microcontrollers #EmbeddedLearning ...

Why Do CPUs Use Registers Instead of Accessing Memory Directly?

When learning assembly or reading compiler output, you will often see instructions like: ADD R0, R1, R2 or LDR R1, [R0] Notice something interesting: Most operations happen between registers, not directly on memory. Why is this the case? The Reason: Speed Registers are located inside the CPU itself. Memory (RAM or Flash) is outside the CPU core and accessed through buses. Because of this, memory access takes significantly longer than register access. Typical comparison: Storage Location Access Speed Register ~1 CPU cycle SRAM multiple cycles Flash even more cycles So if the CPU had to access memory for every operation, programs would run much slower. Example Consider this simple C code: x = a + b; Conceptually, the CPU performs something like: LDR R1, [a] LDR R2, [b] ADD R0, R1, R2 STR R0, [x] The values are first loaded into registers, then the arithmetic operation is performed. Why CPUs Prefer Registers Registers allow the processor to: • execute operations faster • reduce memory acc...

Why Can’t Large Constants Always Fit Inside CPU Instructions?

In assembly, you often see instructions like: ADD R0, R0, #5 Here #5 is embedded directly inside the instruction. This is called Immediate Addressing. But what happens if we write the following C code? x = x + 100000; Instead of placing the value directly inside the instruction, the CPU may generate something like: LDR R1, =100000 ADD R0, R0, R1 Why can’t the processor simply execute: ADD R0, R0, #100000 The Reason: Instruction Size Is Limited Most ARM instructions are 32 bits wide. Those 32 bits must encode several pieces of information: the operation (ADD, SUB, MOV, etc.) the destination register the source register the immediate value Conceptually, the instruction looks like this: [ opcode | registers | immediate value ] Since the instruction must store multiple fields, only a limited number of bits remain for the constant. As a result, very large numbers cannot always fit directly inside the instruction. What Happens When the Constant Is Too Large? When the constant cannot be encod...

Why ARM Keeps PC Ahead?

Image
This behavior simplifies PC-relative addressing. Example instruction: LDR R0, [PC, #0] If the instruction is at: 0x1000 PC already contains: 0x1008 So data is loaded from 0x1008. This is commonly used for literal pools and constant loading. The Key Insight The PC does not point to the instruction currently executing. It points to the instruction already fetched ahead in the pipeline. In ARM state: PC = Current Instruction Address + 8 because the processor pipeline has already fetched future instructions. #EmbeddedSystems #ARM #Firmware #Microcontrollers #EmbeddedLearning #ComputerArchitecture

How a CPU Executes Instructions

Image
Instruction Cycle and Pipelining Explained: Every program you write in C, C++, or assembly eventually becomes a sequence of machine instructions stored in memory. But how does a CPU actually execute those instructions? Inside every processor—from small microcontrollers to powerful server CPUs—the execution of a program happens through a repetitive process known as the Instruction Cycle. Understanding this cycle is one of the most fundamental concepts in computer architecture and embedded systems. What Is the Instruction Cycle? The instruction cycle is the sequence of steps the CPU follows to fetch, understand, and execute an instruction from memory. A CPU does not execute an entire program at once. Instead, it executes one instruction at a time, repeating the same internal process continuously. The classic instruction cycle consists of four stages: Fetch → Decode → Execute → Write Back After completing one instruction, the CPU immediately begins the cycle again for the next...

Why Do Microcontrollers Use Oversampling in UART?

Image
UART communication is asynchronous. That means the transmitter and receiver do not share the same clock. Each device runs using its own oscillator, which can have small timing differences. So how does the receiver know exactly when to read each bit? The answer is oversampling. Instead of sampling the signal once per bit, most microcontrollers sample it 16 times during one bit period. This allows the UART hardware to: • Detect the start bit precisely • Sample near the center of each bit (where the signal is most stable) • Use multiple samples to filter noise • Tolerate small clock mismatches between devices For example, at 9600 baud, a typical UART receiver internally samples at: 9600 × 16 = 153600 samples per second These extra samples help the receiver reconstruct the correct timing of each bit, making communication reliable even when clocks are not perfectly aligned. Without oversampling, UART communication would be much more sensitive to noise and timing errors. Oversampling is one ...

What Happens Inside a Microcontroller in the First Few Microseconds After Power-On?

Image
When power is applied to a microcontroller, the CPU does not immediately start executing your program. Instead, a sequence of hardware-controlled steps occurs to prepare the system for execution. Understanding this startup process helps explain how embedded systems begin running firmware. 1. Power Is Applied When the device receives power: Internal circuits begin initializing. The microcontroller enters a reset state to ensure safe startup. During this stage, the CPU is held in reset and cannot execute instructions yet. 2. Internal Oscillator Starts Most microcontrollers start using an internal RC oscillator. The reason is simple: It starts very quickly It does not require external components It allows the processor to begin operating immediately This oscillator provides the initial system clock. 3. Reset Controller Holds the CPU Even though the clock has started, the CPU is still held in reset by the reset controller. This ensures: Power is stable Clock is stable Internal hardware is ...

Why Do Hardware Timers Run Independently of the CPU?

Image
In microcontrollers, hardware timers operate independently of the CPU. But why? Timers are driven by the peripheral clock, not by CPU instructions. This means the timer continues counting even when the CPU is busy executing other tasks. Because of this design: • The CPU does not need to constantly track time • Timers can generate precise delays • They can trigger periodic interrupts • They enable features like PWM generation and time measurement For example, a timer configured to generate an interrupt every 1 ms will trigger it accurately, regardless of what the CPU is doing. In simple terms: A hardware timer is a dedicated counter inside the microcontroller that runs using its own clock, allowing precise timing without relying on the CPU. ๐Ÿ“š Embedded Systems — WHY Series #EmbeddedSystems #Microcontrollers #Firmware #EmbeddedEngineering #ComputerArchitecture #EmbeddedLearning Timer clock generation

Why Microcontrollers Use Harvard Architecture?

Image
Most microcontrollers use Harvard architecture because it allows faster and more efficient execution. In this architecture: Program memory (Flash) stores instructions Data memory (RAM) stores variables Each uses a separate bus This allows the CPU to: Fetch the next instruction While reading or writing data at the same time. This parallel access improves performance and avoids the Von Neumann bottleneck, where instructions and data share the same bus. That’s why many microcontrollers such as ARM Cortex-M, AVR, and PIC use Harvard architecture #whyseries LinkedIn

why interrupts are better than polling?

Image
Interrupts don’t just save CPU time — they can dramatically extend battery life in embedded systems. Consider a simple polling loop: while(1) {     if(UART_data_available())         read_uart(); } Even when no data arrives, the CPU keeps checking continuously. So the microcontroller stays in active mode, consuming power. With interrupts, the CPU can sleep and wake only when an event occurs. In many microcontrollers: • Active current → few mA • Sleep current → few ยตA That difference is why many battery-powered IoT devices run for months or even years. Interrupts are not just about saving CPU cycles — they are also essential for low-power embedded design. Do you prefer polling or interrupts in your designs? #EmbeddedSystems #Microcontrollers #Firmware #IoT #LowPowerDesign #ARM #TechLearning #why series  LinkedIn

Why most modern microcontrollers are 32-bit

A few years ago, 8-bit microcontrollers like AVR and 8051 were everywhere. Today, most new designs use 32-bit microcontrollers such as ARM Cortex-M. Why did this shift happen? Not because applications suddenly became complex. From my experience working with embedded systems, the shift mostly comes down to cost, performance, and toolchain support. Modern 32-bit microcontrollers offer: • Higher processing power  • Larger address space  • Better compiler optimization  • Advanced peripherals  • Efficient power management What’s interesting is that many 32-bit MCUs today cost almost the same as 8-bit controllers. So designers get more performance without increasing system cost. Examples: • AVR / 8051 → 8-bit  • MSP430 → 16-bit  • ARM Cortex-M → 32-bit STMicroelectronics NXP Semiconductors Microchip Technology Inc. Texas Instruments Infineon Technologies Renesas Electronics Nordic Semiconductor Arm Because of this, 32-bit microcontrollers are now the default choi...

Why RISC is ideal for embedded systems

Image
While revisiting CPU architecture concepts, I summarized why many microcontrollers use RISC architectures. ๐Ÿ”น 1. Predictable execution timing RISC processors typically use fixed-length instructions, making decoding simpler and execution timing more predictable — important for embedded and real-time systems. Example: ADD R1, R2, R3 SUB R4, R5, R6 ๐Ÿ”น 2. Register-based operations RISC follows a Load–Store architecture. ALU operations use registers, while memory is accessed only through LOAD and STORE instructions. Example: LOAD R1, [100] LOAD R2, [104] ADD R3, R1, R2 STORE R3, [108] ๐Ÿ”น 3. Simpler hardware design Because instructions are simple and uniform, the control logic and instruction decoding hardware are simpler, which also helps reduce silicon die size. Example: MOV R1, R2 ADD R3, R1, R4 Conceptually: Register → ALU → Register ๐Ÿ”น 4. Efficient pipelining Fixed instruction formats allow CPUs to pipeline instructions efficiently, improving throughput. Example s...