Semiconductor Memory Architectures: SRAM, DRAM & Non-Volatile
Solid-state digital data storage architectures and device physics: memory organization, word length, bit capacity, 2D/3D matrix addressing, row/column address strobes (RAS/CAS); Static RAM (SRAM) 6-transistor (6T) CMOS bistable cell, precharge lines, read/write stability, and differential sense amplifiers; Dynamic RAM (DRAM) 1-transistor 1-capacitor (1T-1C) trench/stacked cell, destructive readout, capacitive leakage, and periodic refresh scheduling; Non-volatile memory technologies: Mask ROM, fuse PROM, UV-erasable EPROM, floating-gate tunneling EEPROM, and multi-level cell NAND/NOR Flash; SDRAM, DDR protocols, and multi-level cache memory hierarchy.
§7.1 Memory Organization, Matrix Addressing & Decoders
1. General Memory Architecture & Density Classification
A digital semiconductor memory stores binary data in a regular two-dimensional grid of binary memory cells. The overall storage capacity is expressed as:
To access an individual word among $M = 2^k$ addressable locations, the memory requires $k$ binary address input lines ($A_{k-1}, \dots, A_0$). For example, a $64\text{ K} \times 8$ memory chip possesses $2^{16} = 65,536$ words of 8 bits each, requiring $k = 16$ address lines, 8 bidirectional data lines ($D_7 - D_0$), and control lines: Chip Select ($\overline{CS}$), Output Enable ($\overline{OE}$), and Write Enable ($\overline{WE}$).
2. 2D Matrix Addressing & Row/Column Decoders
If $2^k$ words were laid out in a single linear column, the required address decoder would have $2^k$ output lines—requiring an astronomical number of gates ($65,536$ outputs for $k=16$). To achieve high physical density and compact square silicon layouts, memory arrays utilize 2D Matrix Coincident Addressing:
- The $k$ address bits are split into $r$ Row Address bits and $c$ Column Address bits ($k = r + c$).
- The Row Decoder activates exactly one horizontal Word-Line (WL) among $2^r$ rows, simultaneously enabling all $2^c$ storage cells along that row.
- The activated cells place their stored charges onto vertical Bit-Lines (BL).
- The Column Decoder controls a bank of column pass-gates (multiplexers) that route the selected bit-line data to the chip's output buffers.
A $64\text{ K}$-bit array arranged as a $256 \times 256$ matrix requires only one 8-to-256 row decoder and one 8-to-256 column decoder (512 total outputs instead of 65,536).
§7.2 Static RAM (SRAM): The 6T CMOS Memory Cell
1. Structure of the 6-Transistor (6T) CMOS SRAM Cell
Static RAM (SRAM) retains stored data indefinitely as long as DC power is maintained ($V_{DD} > 0$), without requiring periodic refresh cycles. The canonical 6T CMOS SRAM cell comprises:
- Two cross-coupled CMOS inverters ($M_1, M_2, M_3, M_4$) forming a bistable latch with complementary internal storage nodes $Q$ and $\bar{Q}$.
- Two nMOS access pass-transistors ($M_5, M_6$) connecting nodes $Q$ and $\bar{Q}$ to complementary bit-lines ($BL$ and $\overline{BL}$), gated by Word-Line ($WL$).
2. Operational Cycles of the 6T Cell
- Read Cycle:
- Precharge phase: Both $BL$ and $\overline{BL}$ are precharged to $V_{DD}$ (or $V_{DD}/2$) and then floated.
- Assertion phase: Word-line $WL$ is driven HIGH, turning on access transistors $M_5$ and $M_6$.
- Discharge phase: If $Q=0$ and $\bar{Q}=1$, node $Q$ discharges $BL$ through pull-down transistor $M_1$, creating a differential voltage swing $\Delta V = V_{BLB} - V_{BL} \approx 100 - 200\text{ mV}$.
- Sensing phase: A sensitive analog differential sense amplifier strobes, detecting $\Delta V$ and amplifying it rapidly to full CMOS logic levels ($0\text{ V}$ or $V_{DD}$).
- Write Cycle:
- Strong write drivers overdrive $BL$ and $\overline{BL}$ to opposite supply rails (e.g., $BL = 0\text{ V}, \overline{BL} = V_{DD}$ to write a 0).
- $WL$ is asserted HIGH. The strong pull-down on $BL$ overpowers the weaker internal pMOS pull-up transistor, flipping the cross-coupled latch into the new state.
Cell Sizing (Read Stability & Write Margin): To prevent the cell from accidentally flipping its state during a read operation (read disturbance), pull-down transistors $M_1, M_3$ must be made stronger than access transistors $M_5, M_6$ ($\beta_{\text{pull-down}} / \beta_{\text{access}} \ge 1.2 - 1.5$). Conversely, to ensure data can be written successfully, access transistors must be stronger than pull-up transistors $M_2, M_4$ ($\beta_{\text{access}} / \beta_{\text{pull-up}} \ge 1.0$).
§7.3 Dynamic RAM (DRAM): The 1T-1C Cell & Refresh Scheduling
1. The 1-Transistor 1-Capacitor (1T-1C) DRAM Cell
Robert Dennard (IBM, 1968) patented the 1T-1C DRAM cell, which slashed silicon area from 6 transistors down to a single access transistor $M$ and an integrated storage capacitor $C_s$ ($C_s \approx 25 - 35\text{ fF}$):
- Logic '1' is stored as a packet of charge on $C_s$ ($V_C \approx V_{DD}$).
- Logic '0' is stored as discharged state ($V_C \approx 0\text{ V}$).
Because the cell area is minuscule ($4F^2 - 6F^2$, where $F$ is the lithographic feature size), DRAM achieves gigabit storage densities orders of magnitude higher than SRAM, making it the universal choice for computer main system memory.
2. Charge Sharing & Destructive Readout
When Word-Line $WL$ is asserted during a read operation, the storage capacitor $C_s$ shares its charge with the much larger parasitic capacitance of the long bit-line $C_{BL}$ ($C_{BL} \sim 100 - 300\text{ fF} \approx 10 C_s$):
Precharging the bit-line to $V_{DD}/2$ produces a minute voltage perturbation:
A cross-coupled regenerative sense amplifier detects $\Delta V$ and swings the bit-line fully to $V_{DD}$ or $0\text{ V}$. Because charge sharing partially discharges $C_s$, the readout is destructive. The sense amplifier must immediately rewrite the amplified logic level back into $C_s$ before closing the word-line (Restore Cycle).
3. Capacitor Leakage & Periodic Refresh Scheduling
Due to subthreshold MOSFET leakage and reverse-biased $p$-$n$ junction leakage currents, charge leaks off $C_s$ with a time constant of tens of milliseconds. To prevent catastrophic data loss, every row of the DRAM array must be read and rewritten (refreshed) periodically (typically every $64\text{ ms}$ at $85^\circ\text{C}$):
Modern DRAM controllers interleave Distributed Auto-Refresh or Self-Refresh commands between normal read/write cycles, consuming less than $1 - 2\%$ of total memory bus bandwidth.
§7.4 Non-Volatile Memories: ROM, PROM, EPROM, EEPROM & Flash
1. Read-Only Memory (ROM) Evolution
Non-volatile semiconductor memories retain stored data indefinitely without requiring power supplies:
- Mask ROM: Data is permanently hardwired during wafer fabrication using a custom photolithographic contact mask. Highest density and lowest cost per bit in mass production, but zero programmability.
- Programmable ROM (PROM): Fabricated with microscopic Nichrome or polycrystalline silicon fuses in series with each memory cell. Programmed once by blowing selected fuses using high-current pulses (One-Time Programmable - OTP).
- Erasable PROM (UV-EPROM): Utilizes a Floating-Gate MOSFET (FGMOS) with a completely isolated conductive polysilicon gate embedded inside silicon dioxide dielectric. High-voltage pulses ($V_{PP} \approx 12 - 21\text{ V}$) inject electrons onto the floating gate via Hot-Carrier Injection (HCI). Stored electrons shift the transistor's threshold voltage ($V_t$), programming it to state 0. Erased by shining ultraviolet light ($254\text{ nm}$) through a quartz window on the chip package for 20 minutes, exciting electrons over the $\text{SiO}_2$ potential barrier.
- Electrically Erasable PROM (EEPROM): Employs ultra-thin tunnel oxide ($d_{\text{ox}} < 10\text{ nm}$) beneath the floating gate, enabling bidirectional electrical erasure via Fowler-Nordheim (F-N) Quantum Mechanical Tunneling. Can be erased and reprogrammed byte-by-byte in circuit.
2. Modern Flash Memory: NAND vs NOR Architectures
Invented by Fujio Masuoka (Toshiba, 1984), Flash memory erases blocks of cells simultaneously in a single flash operation via Fowler-Nordheim tunneling:
| Attribute | NOR Flash | NAND Flash |
|---|---|---|
| Cell Interconnection | Parallel (like NOR gate) | Series strings of 32 to 128 cells (like NAND gate) |
| Random Access Speed | Very Fast ($50 - 80\text{ ns}$) | Slow initial access ($25\text{ \mu s}$) |
| Serial Throughput | Moderate | Extremely High ($> 1\text{ GB/s}$) |
| Cell Area Density | Large ($10F^2$) | Extremely Compact ($4F^2$, 3D vertical stacked $> 200$ layers) |
| Primary Application | BIOS, Router Firmware (eXecute-In-Place) | Solid-State Drives (SSDs), USB drives, Smartphones |
§7.5 Advanced Memory Architectures: SDRAM, DDR & Cache Hierarchy
1. Synchronous DRAM (SDRAM) & Double Data Rate (DDR)
Early asynchronous DRAMs required address strobe handshakes ($\overline{RAS}, \overline{CAS}$) that bottlenecked high-speed microprocessors. Synchronous DRAM (SDRAM) synchronizes all control, address, and data lines to the master CPU system clock using internal pipelining and multi-bank prefetching.
Double Data Rate (DDR) SDRAM: Transfers data on both the rising and falling edges of each clock cycle, doubling throughput at identical clock frequencies:
- DDR1: 2-bit prefetch buffer (2 data words per clock cycle).
- DDR2: 4-bit prefetch buffer.
- DDR3: 8-bit prefetch buffer.
- DDR4: 16-bit prefetch buffer with independent memory bank groups.
- DDR5: Dual 32-bit channels per module, on-die ECC, data transfer rates exceeding $6400\text{ MT/s}$.
2. Multi-Level Cache Memory Hierarchy
Because processor execution speed ($3 - 5\text{ GHz}$, sub-nanosecond cycles) is orders of magnitude faster than DRAM latency ($50 - 70\text{ ns}$), modern computer architecture incorporates a multi-tiered cache hierarchy exploiting the Principle of Locality (Temporal and Spatial):
The Average Memory Access Time (AMAT) is:
Design a $64\text{ KB}$ ($65,536 \text{ bytes}$) microprocessor memory subsystem using standard $16\text{ KB} \times 8$ SRAM memory chips. (a) Determine the number of memory chips required. (b) Calculate the total number of address bus lines required for the subsystem and how many address lines connect directly to each chip. (c) Design the chip-select address decoding circuit using a 2-to-4 binary decoder and specify the hexadecimal address range for each memory block.
Four 16 KB x 8 chips are required to synthesize 64 KB.
Address lines A0 through A13 connect in parallel to all 4 chips to select a byte within each chip.
A 2-to-4 decoder decodes A15 and A14 to enable one of the four chips.
List continuous address space from 0000H to FFFFH.
N_{\text{chips}} = 4; \quad 16 \text{ total address lines } (A_{13}-A_0 \text{ to chips}, \ A_{15}-A_{14} \text{ to 2:4 decoder}); \quad \text{Range: } 0000_{16} - \text{FFFF}_{16}
An $8\text{ Gb}$ DDR4 DRAM memory chip is organized internally into $8192$ rows. All $8192$ rows must be refreshed within a maximum retention window of $t_{\text{ret}} = 64.0\text{ ms}$. Each individual row refresh cycle takes $t_{RC} = 45.0\text{ ns}$. (a) Calculate the average refresh frequency (how often a refresh command must be issued). (b) Calculate the total time consumed per $64\text{ ms}$ interval solely for refreshing. (c) Determine the percentage of memory bandwidth lost to refresh operations.
A refresh command must be executed every 7.81 microseconds.
Multiply number of rows by row cycle time.
Refresh consumes only 0.58% of total memory bandwidth.
t_{\text{REFI}} = 7.81\text{ \mu s}, \quad t_{\text{total}} = 368.6\text{ \mu s per } 64\text{ ms}, \quad \text{Bandwidth Overhead} = 0.576\%
A modern computer architecture features a two-level cache hierarchy: Level 1 ($L_1$) cache has an access time of $t_1 = 1.20\text{ ns}$ and a hit rate of $H_1 = 94.0\%$; Level 2 ($L_2$) cache has an access time of $t_2 = 6.00\text{ ns}$ and a local hit rate of $H_2 = 85.0\%$; Main DRAM memory has an access latency of $t_{MM} = 55.0\text{ ns}$. (a) Calculate the global miss rate for the cache system. (b) Calculate the Average Memory Access Time (AMAT) of the processor. (c) By what factor would memory performance degrade if the caches were eliminated?
Only 0.9% of memory requests reach main DRAM.
Formulate multi-level AMAT equation.
Average access time is only 2.06 ns.
Without caches, memory access would be nearly 27 times slower!
\text{Global Miss Rate} = 0.90\%, \quad \text{AMAT} = 2.055\text{ ns}, \quad \text{Speedup over Raw DRAM: } 26.8\times
Solved University Examination Problems
Step-by-step mathematical solutions to classic university honors examination questions.