                            C Memory Models

                            Matthew Probert
                           Servile  Software


When programming in C on the DOS platform (normal PC), the programmer 
has available a range of memory models: Tiny, Small, Medium, Compact, 
Large and Huge. Each memory model is defined by the way code, and data 
is addressed. Basically there are three types of addressing:
 
        near    16 bit address within a single 64K segment
        far     32 bit address within a 1mB range
        huge    32 bit address with normalised pointers

The far and huge addressing modes are most significant when accessing 
data. A 32 bit standard pointer is comprised of a 16 bit segment and 
16 bit offset (although only 4 bits of the segment are used). When 
arithmetic is applied to a 32 bit pointer the offset is changed, but 
not the segment. This results in wrap-around occuring, and only a 
single segment (64K) being readily addressable. A normalised pointer
is a 20 bit pointer which when arithemtic is applied to it is changed 
as a single entity, thereby allowing it to address an entire 1mB 
range.

Each C program is comprised of a number of memory classes: code, 
data, stack, heap and farheap. 

The code class is where the program instructions are stored. 

The data class is where all global and static variables are stored. 

The stack class provides a dynamic and temporary area for the storage 
of parameters and automatic variables. 

The heap is dynamic memory allocated with malloc() class functions. 
The farheap is the remainder of memory which may or may not be the 
heap, depending upon the memory model. 

The Tiny memory model places all memory classes (except the far heap) 
into a single segment. 16 bit pointers are used for functions and 
variables.

The Small memory model uses two seperate segments. One for code and 
one for data. 16 bit pointers are used for both code and data 
addressing.

The Medium memory model uses 32 bit pointers for addressing code, and 
allows multiple code segments.

The Compact memory model uses 16 bit pointers for addressing code 
and 32 bit pointers for addressing data, allowing a single code 
segment and multiple data segments. In addition the compact memory 
model puts the stack in a segment of its own, allowing a full 64K 
stack segment. 

The Large memory model uses 32 bit pointers for addressing both code 
and data, allowing multiple code segments and multiple data segments. 
In addition the large memory model puts the stack in a segment of its 
own, allowing a full 64K stack segment.

The Huge memory model uses normalised 32 bit pointers for addressing 
both code and data, allowing multiple code segments and multiple data 
segments. In addition the huge memory model puts the stack in a 
segment of its own, allowing a full 64K stack segment. Unlike the 
compact or large memory model, the huge memory model allows more than 
64K of static data to be declared.

There are reasons for having different memory models. 16 bit address 
pointers compile into smaller and faster code than 32 bit address 
pointers, and these in turn compile into smaller and faster code than 
normalised (huge) address pointers. The advice being, use the smallest 
memory model you can!

I have mentioned multiple code and data segments, but how do you 
declare them? By using multiple source modules is the answer.

When a program is compiled in a large code model (Medium, Large, Huge) 
each source module compiles into a separate code segment. Together 
these code segments may total more than 64K, individually (with the 
exception of the huge memory model) they must each be less than 64K 
when compiled.

Too much global data?

If you're declaring too much global data as in the case of:

/*
    This will cause a compiler error
*/
char arr1[65535];
char arr2[65535];

try declaring the variables as FAR such as:

/*
    Borland C is quite happy with this
*/
char far arr1[65535];
char far arr2[65535];

The Borland C compilers place each 'far' data declaration into its own 
data segment, allowing (depending upon memory model) a program to 
declare in excess of 64K of global data.

Speeding Up Large Code Memory Models:

Okay, you've written the mother of all C programs. It is thousands of 
lines of code and is a number of source modules and is compiled in the 
Medium or Large memory model. How do you speed it up? If there are 
functions which are called only by a few other functions, place all 
those functions into a single source module and declare them as static 
and near. By declaring a function as static, you ensure that no other 
function can attempt to call it. By declaring a function as near you 
force the compiler to override the memory model and compile 16 bit 
calling and return code for that function. This not only reduces the 
size of the compiled program, but also makes it execute faster. 
Remember though to prototype these functions within the source module, 
otherwise the compiler may generate far calls to them!
