
                                                                 
                   Code and Concept: How to Gracefully Load      
                       and Unload a Multi-Threaded NLM           
                                                                 
             

                                  Matt Hagen
                                  Consultant
                         Systems Engineering Division

               ͻ
                     PgDn to Scroll or follow link for      
                           Table of Contents               
               ͼ

                                  Abstract:

     This AppNote explains how to build a multi-threaded NLM that loads and
     unloads correctly.



                                   Contents

           Introduction
           Processes and Threads
           Conservation of Resources
           Multi-threaded Scenario
           A Graceful Load
           A Graceful Unload
           Example Program
           Notes


     Disclaimer

     Novell, Inc. makes no representations or warranties with respect to the
     contents or use of these Application Notes (AppNotes) or of any of the
     third-party products discussed in the AppNotes. Novell reserves the right
     to revise these AppNotes and to make changes in their content at any
     time, without obligation to notify any person or entity of such revisions
     or changes. These AppNotes do not constitute an endorsement of the
     third-party product or products that were tested. Configuration(s) tested
     or described may or may not be the only available solution. Any test is
     not a determination of product quality or correctness, nor does it ensure
     compliance with any federal, state or local requirements. Novell does not
     warranty products except as stated in applicable Novell product
     warranties or license agreements.

     Copyright (c) 1991 by Novell, Inc., Provo, Utah. All rights reserved.

     As a means of promoting NetWare AppNotes, Novell grants you without
     charge the right to reproduce, distribute and use copies of the AppNotes,
     provided you do not receive any payment, commercial benefit or other
     consideration for the reproduction or distribution, or change any
     copyright notices appearing on or in the document.

Introduction

     Frequently, complex NLMs need to create (or spawn) several threads. In
     turn, each thread usually allocates one or more NetWare resources. For
     example, a job server may need to create a pool of threads, each ready to
     service a request from a network client. Such multi-threaded NLMs require
     careful synchronization between threads especially during load and unload
     time. This Application Note explains how to use CLIB routines like
     ThreadSwitch(), WaitOnLocalSemaphore, and SignalLocalSemaphore to
     initialize such an NLM. It also explains how to use the CLIB routine
     signal() to dismantle a multi-threaded NLM. In the process, the AppNote
     explores the concept of a NetWare thread and explains the Principle of
     Thread Resource Responsibility. I've also included an example program,
     GRACE.C, that runs on top of NetWare v3.1 with CLIB v3.1 and STREAMS
     v3.10 loaded. The program compiles and links using C Network Compiler/386
     v1.10.

     Before examining the initialization and deinitialization of a
     multi-threaded NLM, let's define the term thread and then discuss the
     reponsibility of a thread with regard to resource allocation and
     deallocation.

Processes and Threads

     A NetWare process is defined by a structure called a Process Control
     Block (PCB). The values in this structure determine the process' name,
     priority, ESP (stack pointer), EIP (code pointer), stack size, and other
     attributes. A NetWare thread is the CLIB version of a NetWare process.
     The thread structure includes a PCB structure along with other CLIB-
     specific information. For our purposes, the terms process and thread are
     interchangeable.

     The NetWare operating system creates a handful of processes at boot time.
     NLMs can also create one or more processes. Each process has a certain
     job to do. For example, the NetWare Command Line Process periodically
     checks the server's console command line to see if the console operator
     has typed LOAD, NAME, or some other console command. If the process finds
     a console command on the command line, it executes the command; otherwise
     it does nothing. The NetWare Cache Update Process periodically checks
     server RAM to see if a block of memory needs to be written to disk. An
     NLM might create a process that polls a job queue.

     A process can do its job only when it has control of the host machine's
     CPU. This state is called the running state. A process can also be
     waiting or sleeping.  Figure 1 illustrates these three states.

     The running process has control of the machine's CPU. It is the only
     process (in a single-CPU machine) that is currently accomplishing its
     job. It can be interrupted only by a hardware interrupt. No other process
     can force it to relinquish control of the CPU.

     A waiting process resides on the run queue, a linked list of PCBs,
     ordered by priority, each waiting to become the running process. These
     processes are considered awake.

     A sleeping process (by definition) is not the running process and is not
     linked into the run queue. There is no sleep queue in the NetWare
     operating system. A sleeping process may be linked into one of several
     dissimilar lists. One process may be asleep waiting for keyboard input
     (for example, a process in INSTALL.NLM). Another process may be asleep
     waiting for a disk write to complete (for example, the cache update
     process mentioned above). Still another process may be asleep waiting to
     service a request packet (for example, a file service process). A
     sleeping process cannot reawaken itself. The reawakening must be done by
     another process or by a hardware interrupt.

     A process can experience any of the following four state changes:

         Run queue to running process. The process, after working its way to
          the head of the run queue, gains control of the CPU.

         Running process to run queue. The process relinquishes control of
          the CPU but does not go to sleep. Instead, it relinks itself into
          the run queue, usually at the end of its priority. The CLIB function
          ThreadSwitch() effects this change of state.

         Running process to sleep. The process relinquishes control and goes
          to sleep (goes somewhere other than the run queue). Before going to
          sleep, a process must make the pointer to its PCB available to some
          other process or interrupt that will eventually reawaken the
          process.

         Sleep to run queue. Some other process or interrupt relinks the
          awakening process into the run queue.

     A context switch occurs when one process relinquishes control of the CPU
     and another process gains control. The relinquishing process must save
     its EIP and ESP pointers (on its own stack) and initialize the EIP and
     ESP registers for the new process before actually giving up control of
     the CPU.

     Now that we know what a NetWare process is, let's explore the
     responsibilities of a process regarding resources.

Conservation of Resources

     Frequently, a process created by an NLM allocates one or more resources.
     The term resource almost always refers to (1) an unorganized block of
     memory or (2) a structure (a block of memory designated for a specific
     use). Examples of resources include memory, sockets, semaphores, ECBs,
     screens, connections, tasks, and interrupt vectors. Typically, a process
     allocates resources just after it is created, uses the resources during
     its lifetime (which may include all the state changes mentioned above),
     and returns the resources to NetWare just before it is destroyed. This
     classic thread scenario follows a concept that I call the Principle of
     Thread Resource Responsibility which states that, in general, the thread
     that allocates a resource should deallocate the resource. Adherence to
     this principle eliminates the need for an NLM to maintain global lists of
     resources, which tends to make the NLM less efficient and unnecessarily
     complicated.

     This principle also requires multi-threaded NLMs to carefully orchestrate
     the creation and destruction of the threads.  Let's look at an example.

Multi-threaded Scenario

     Assume that you want to create an NLM that listens for requests from
     network clients on an IPX socket and services the requests. The NLM must
     be able to service several requests at the same time. It must also
     allocate a screen and display certain statistics on the screen, updating
     the values on the screen every second.

     One approach is to create an NLM that consists of three thread groups.
     The first thread group includes the main thread (the one created by CLIB)
     that listens for request packets. The second group includes a pool of
     five worker threads that service the requests. The third group includes
     the update thread that refreshes the counters on the NLM's screen every
     second.

     The life of each thread in this example scenario consists of an
     initialization phase, a run phase, and a deinitialization phase. Let's
     explore these three phases for each type of thread.

     Main Thread. During initialization, the main thread allocates a socket, a
     semaphore, and a pool of listen ECBs. It also posts the listen ECBs,
     creates the other threads, and enters the run phase of its life by going
     to sleep on the semaphore.

     When an incoming packet arrives, the interrupt service routine (ISR) of
     the LAN card driver bumps the main thread from the semaphore to the run
     queue. After ascending the run queue and becoming the running process,
     the main thread wakes up one of the worker threads, hands the request to
     the worker (in the form of an ECB), and goes back to sleep on the
     semaphore. The main thread does this repeatedly during the run phase of
     its lifetime.

     During deinitialization, the main thread wakes up the worker threads and
     the update thread and allows them to return any NetWare resources they
     allocated. Then it deallocates its own resources (the listen ECBs, the
     semaphore, and the socket).

     Worker Thread. During initialization, each worker thread allocates a
     socket, a pool of listen ECBs, and a pool of send ECBs. It then posts the
     listen ECBs and enters the run phase of its life by going to sleep on a
     global list of sleeping worker threads.

     When the main thread receives an incoming request packet, the main thread
     unlinks the next available worker thread from the global list, hands the
     request to the worker (in the form of an ECB), and kicks the worker
     thread awake by linking it into the run queue. Once the worker thread
     becomes the running process, it services the request, alternately hopping
     between the running process position and the run queue. Once the request
     is serviced, the worker thread relinks itself into the global list of
     sleeping threads and goes to sleep.

     During deinitialization, the worker thread deallocates the send ECBs, the
     receive ECBs, and the socket.

     Update Thread. During initialization, the update thread allocates a
     screen, writes a title on the screen, and quickly enters the run phase of
     its life.

     During the run phase of its lifetime, the update process displays the
     values of certain counters on the screen, puts itself on a list
     maintained by the timer interrupt, and goes to sleep. After one second,
     the timer interrupt wakes up the update process by linking it into the
     run queue. The update process continues this loop indefinitely.

     During deinitialization, the update process returns the screen.

     The challenge involved in building an multi-threaded NLM like the one
     described above is two-fold:  (1) the NLM must allow each thread to
     complete initialization before it (the NLM) admits to being fully
     initialized, and (2) the NLM must allow each thread to complete
     deinitialization before the NLM allows itself to be unloaded. Let's
     explore the specifics of initialization and deinitialization further.

A Graceful Load

     This section presents two techniques that allow multi-threaded NLMs to
     load gracefully. The first technique requires that the main thread poll a
     global variable until a newly created child thread succeeds (or fails) to
     initialize. The example code below illustrates this technique. In the
     example, the main thread (created by CLIB) sets threadInitFlag to WAIT.
     Next, it calls BeginThreadGroup() to create the update thread. Then it
     jumps back and forth between the run queue and the CPU, polling
     threadInitFlag and waiting for the update thread to set to flag to
     CONTINUE or EXIT. Meanwhile, the update thread attempts to allocate a
     resource (a screen). If it fails, it sets threadInitFlag to EXIT and
     exits. Otherwise, it sets the flag to CONTINUE and then continues with
     its job, eventually going to sleep. When the main thread discovers that
     threadInitFlag no longer equals WAIT, it stops polling the flag and acts
     according to the new value. If the flag is set to EXIT, the main thread
     aborts the LOAD attempt. If the flag is set to CONTINUE, the main thread
     continues to load the NLM. Here is the example code:

     #define CONTINUE 0                 /* values for the global variable */
     #define WAIT 1
     #define EXIT 2

     BYTE threadInitFlag;               /* global initialization variable */

/**************************************************************************
* main
**************************************************************************/

     main(
     {
     ...

          threadInitFlag=WAIT;
     if(BeginThreadGroup(ScreenRoutine,NULL,NULL,NULL)==EFAILURE)
     {
          ConsolePrintf("  Cannot create screen thread.\n");
          ...
          goto localExit;
     }

     while(threadInitFlag==WAIT) /* wait for the new thread to succeed
     or fail */
          ThreadSwitch();

          if(threadInitFlag!=CONTINUE)    /* if the new thread failed, exit */
     {
     ...
     goto localExit;
     }

          ...                              /* otherwise continue */

     localExit:;
     }

/**************************************************************************
* ScreenRoutine
**************************************************************************/

     void ScreenRoutine(
     void)
     {
     int screen;

          screen=CreateScreen("Graceful Screen",AUTO_DESTROY_SCREEN);
     if(screen==EFAILURE)
     {
          ConsolePrintf("  Cannot create screen.\n");
          threadInitFlag=EXIT;        /* if can't allocate resource, set
     flag */
          goto localExit;
     }

          threadInitFlag=CONTINUE;         /* after all resources allocated,
     set flag */

          ...                              /* eventually sleep */

     localExit:;
     }

     Note that this technique employs polling. As a general rule of thumb,
     NLMs should avoid polling because it violates the Principle of Minimal
     CPU Utilization which states that a thread should endeavor to use the CPU
     as few times as possible for the shortest intervals possible while still
     accomplishing its task. In the example above, however, the update thread
     does not relinquish control while allocating the screen, which guarantees
     that the main thread will have to poll the flag only once. In this case,
     polling is the best choice because it is simple and it does not degrade
     server performance.

     The second technique requires two global variables:  the main thread's
     thread ID and a flag. The example code below illustrates this technique.
     First, the main thread calls GetThreadID() to initialize mainThread and
     sets threadInitFlag to EXIT. Next, the main thread calls
     BeginThreadGroup() to create the update thread. Then, the main thread
     calls SuspendThread() to go to sleep, allowing the update thread to run.
     If the update thread allocates the screen successfully, it sets
     threadInitFlag to CONTINUE and continues, eventually going to sleep.
     Otherwise, it leaves the flag set to EXIT and exits. In either case,
     after attempting to allocate the resource, it calls ResumeThread() to
     wake up the main thread. When the main thread regains the CPU, it checks
     the flag and acts accordingly. If the flag is still set to EXIT, the main
     thread aborts the LOAD attempt. If is set to CONTINUE, it continues to
     load the NLM. Here is the example:

     #define CONTINUE 0                    /* values for the global variable
     */
     #define EXIT 1

     int mainThread;
     BYTE threadInitFlag;                  /* global initialization variable
     */

/**************************************************************************
* main
**************************************************************************/

     main(
     {
     ...

          mainThread=GetThreadID();

          threadInitFlag=EXIT;
     if(BeginThreadGroup(ScreenRoutine,NULL,NULL,NULL)==EFAILURE)
     {
          ConsolePrintf("  Cannot create screen thread.\n");
          ...
          goto localExit;
     }

          SuspendThread(mainThread);
     if(threadInitFlag==EXIT)         /* update thread failed to
     initialize */
     goto localExit;                  /* refuse to load NLM */

          ...                              /* eventually sleep */

     localExit:;
     }

/***************************************************************************
ScreenRoutine
**************************************************************************/

     void ScreenRoutine(
     void)
     {
     int screen;

          screen=CreateScreen("Graceful Screen",AUTO_DESTROY_SCREEN);
     if(screen==EFAILURE)
     {
          ConsolePrintf("  Cannot create screen.\n");
          ResumeThread(mainThread);   /* leave the flag set to EXIT */
          goto localExit;
     }

          threadInitFlag=CONTINUE;         /* after all resources allocated,
     set flag */
     ResumeThread(mainThread);

          ...

     localExit:;
     }

     This technique is useful when the child thread may need to go to sleep
     while allocating resources.

A Graceful Unload

     Assume that a console operator has just entered UNLOAD at the console to
     unload the NLM described in "Multi-threaded Scenario" above. Chances are
     all threads in all three thread groups are asleep:  the main thread on
     the semaphore, the worker threads on the global sleep list, and the
     update thread on the timer interrupt list. When the NetWare Command Line
     Process (mentioned earlier in this AppNote) becomes the running process,
     it finds UNLOAD on the command line and begins to unload the NLM. By
     default, CLIB will allow the Command Line Process to kill all the threads
     in the NLM without letting them run again. This would be disastrous in
     our case, since all the threads have resources that they still need to
     return to NetWare. For such cases, CLIB provides a way for threads to
     avoid the lethal stroke of the Command Line Process' axe long enough to
     return resources. This method is in the form of a routine called
     signal(). When a thread calls signal() like this,

          signal(SIGTERM,SignalRoutine);

     the Command Line Process, at unload time, will not kill the thread until
     after the Command Line Process has called and returned from the
     SignalRoutine(), which is defined by the NLM.

     The following example shows how an NLM can use signal() and a few global
     variables to orchestrate a graceful unload:

     #define CONTINUE 0
     #define EXIT 2

     LONG semaphore;
     BYTE threadCount=0;
     BYTE globalExitFlag=CONTINUE;

/***************************************************************************
main
**************************************************************************/

     main(
     {
     signal(SIGTERM,SignalRoutine);   /* let me run at unload time */
     threadCount++;                   /* add self to thread count */

          /* allocate semaphore */
     /* create update process */

          while(TRUE)
     {
          WaitOnLocalSemaphore(semaphore); /* sleep on semaphore */
          if(globalExitFlag!=CONTINUE)     /* check for unload */
          goto exit2;                      /* yes, unload */

               /* do work */                    /* no, do not unload */
     }

     exit2:;
     ResumeThread(updateThread);

     exit1:;
     while(threadCount>1)               /* wait until all other threads
                                        deinitialize */
          ThreadSwitch();

     exit0:;
     /* close semaphore */

          threadCount#--;                  /* subtract self from thread count
          */
     }

/***************************************************************************
SignalRoutine
**************************************************************************/

     void SignalRoutine(
     void)
     {
     globalExitFlag=EXIT;

          SignalLocalSemaphore(semaphore);   /* wake up the main thread so
                                             that it */
                                        /* can wake up all the other threads
                                        */
     while(threadCount>0)               /* wait until all threads, including
                                        */
     ThreadSwitch();                  /* the main thread, */
     /* fully deinitialize */
     }

/***************************************************************************
ScreenRoutine
**************************************************************************/

     void ScreenRoutine(
     void)
     {
     signal(SIGTERM,SignalRoutine);   /* let me run at unload time */
     threadCount++;                   /* add self to thread count */

          /* create screen */

          while(TRUE)
     {
          /* update screen */

               delay(1000);

               if(globalExitFlag!=CONTINUE) /* check for unload */
               goto exit1;             /* yes, unload */
     }

     exit1:;
     /* destroy screen */             /* give back resource */

     exit0:;
     threadCount--;                   /* subtract self from thread count */

          SuspendThread(screenThread);     /* wait to be killed */
     }

     At unload, the Command Line Process calls the SignalRoutine() before
     killing any of the threads that called signal(). SignalRoutine() wakes up
     the main thread and then refuses to return to the Command Line Process
     until all threads have fully deinitialized. Once awake, the main thread
     wakes up all the other threads (in this case only the update thread) and
     allows them to return their resources. Then the main thread returns its
     resources and decrements threadCount, allowing SignalRoutine() to return.
     This, in turn, allows the Command Line Process to continue its job of
     killing the sleeping threads and unloading the NLM.

Example Program

     This section presents GRACE.C, the source code for an example NLM that
     reflects the important features of the NLM described in "Multi-threaded
     Scenario" above. The main thread creates five worker threads and one
     update thread. Each thread allocates enough resources to make it
     necessary to be careful during initialization and deinitialization.
     During the life of the NLM, the main thread sleeps on the semaphore
     (presumably waiting for request packets), the worker threads sleep on a
     global list (waiting to process request packets), and the update thread
     sleeps on the timer interrupt list, waking up once every second to update
     the screen. The example shows how to gracefully initialize and
     deinitialize a multi-threaded NLM.

/***************************************************************************
GRACE.C
**************************************************************************\

     #include <nwtypes.h>
     #include <nwsemaph.h>
     #include <nwipxspx.h>
     #include <conio.h>
     #include <process.h>
     #include <stdio.h>
     #include <stdlib.h>
     #include <signal.h>
     #include <errno.h>

     typedef struct WorkerStructure
     {
     struct WorkerStructure *Slink;
     int threadID;
     }WORKER;

     void SignalRoutine(
     void);

     void WorkerRoutine(
     void);

     void UpdateRoutine(
     void);

     #define CONTINUE 0
     #define WAIT 1
     #define EXIT 2

     LONG semaphore;
     BYTE threadCount=0;
     WORKER *sList=NULL;
     int updateThread;
     BYTE globalExitFlag=CONTINUE;
     BYTE threadInitFlag;

/***************************************************************************
main
**************************************************************************/

     main(
     int argc,
     char *argv[])
     {
     int a;
     WORKER *w;

          signal(SIGTERM,SignalRoutine);
     threadCount++;

          semaphore=OpenLocalSemaphore(0);
     if(semaphore==NULL)
     {
          ConsolePrintf("  Cannot allocate semaphore.\n");
          globalExitFlag=EXIT;
          goto exit0;
     }

          for(a=0;a<5;a++)
     {
          threadInitFlag=WAIT;
          if(BeginThreadGroup(WorkerRoutine,NULL,NULL,NULL)==EFAILURE)
          {
               ConsolePrintf("  Cannot create session thread #%d.\n",a+1);
               globalExitFlag=EXIT;
               goto exit1;
          }

               while(threadInitFlag==WAIT)
               ThreadSwitch();

               if(threadInitFlag!=CONTINUE)
          {
               globalExitFlag=EXIT;
               goto exit1;
          }
     }

          threadInitFlag=WAIT;
     if(BeginThreadGroup(UpdateRoutine,NULL,NULL,NULL)==EFAILURE)
     {
          ConsolePrintf("  Cannot create screen thread.\n");
          globalExitFlag=EXIT;
          goto exit1;
     }

          while(threadInitFlag==WAIT)
          ThreadSwitch();

          if(threadInitFlag!=CONTINUE)
     {
          globalExitFlag=EXIT;
          goto exit1;
     }

          while(TRUE)
     {
          WaitOnLocalSemaphore(semaphore);
          if(globalExitFlag!=CONTINUE)
               goto exit2;

                    /* take worker off sleep list */
               /* give request to worker */
               /* wake up worker */
     } exit2:;
     ResumeThread(updateThread);

     exit1:;
     w=Slist;
     while(w!=NULL)
     {
          ResumeThread(w->threadID);
          w=w->Slink;
     }

          while(threadCount>1)
          ThreadSwitch();

     exit0:;
     CloseLocalSemaphore(semaphore);

          threadCount--;
     }

/***************************************************************************
SignalRoutine
**************************************************************************/

     void SignalRoutine(
     void)
     {
     globalExitFlag=EXIT;

          SignalLocalSemaphore(semaphore);

          while(threadCount>0)
          ThreadSwitch();
     }

/***************************************************************************
WorkerRoutine
**************************************************************************/

     void WorkerRoutine(
     void)
     {
     WORKER *w;
     void *ecb;

          signal(SIGTERM,SignalRoutine);
     threadCount++;

          w=malloc(sizeof(WORKER));
     if(w==NULL)
     {
          ConsolePrintf("  Cannot allocate WORKER structure.\n");
          threadInitFlag=EXIT;
          goto exit0;
     }

          ecb=malloc(sizeof(IPX_ECB));
     if(ecb==NULL)
     {
          ConsolePrintf("  Cannot allocate session memory.\n");
          threadInitFlag=EXIT;
          goto exit1;
     }

          w->threadID=GetThreadID();
     threadInitFlag=CONTINUE;

          while(TRUE)
     {
          w->Slink=Slist;
          Slist=w;
          SuspendThread(w->threadID);

               if(globalExitFlag!=CONTINUE)
               goto exit2;

               /* service the request */
     }

     exit2:;
     free(ecb);

     exit1:;
     free(w);

     exit0:;
     threadCount--;
     SuspendThread(GetThreadID());
     }

/***************************************************************************
UpdateRoutine
**************************************************************************/

     void UpdateRoutine(
     void)
     {
     int screen;
     LONG count=0;

          signal(SIGTERM,SignalRoutine);
     threadCount++;

          updateThread=GetThreadID();

          screen=CreateScreen("Graceful Screen",AUTO_DESTROY_SCREEN);
     if(screen==EFAILURE)
     {
          ConsolePrintf("  Cannot create screen.\n");
          threadInitFlag=EXIT;
          goto exit0;
     }

          HideInputCursor();

          if(DisplayScreen(screen)!=ESUCCESS)
     {
          ConsolePrintf("  Cannot display screen.\n");
          threadInitFlag=EXIT;
          goto exit1;
     }

          threadInitFlag=CONTINUE;

          gotoxy(20,0);
     printf("Graceful Init/Deinit Example Screen");

          while(TRUE)
     {
          SetCurrentScreen(screen);

               gotoxy(0,2);
          printf("Update Count = %u\n",count++);
          printf("Thread Count = %u",threadCount);
          delay(1000);

               if(globalExitFlag!=CONTINUE)
               goto exit1;
     }

     exit1:;
     DestroyScreen(screen);

     exit0:;
     threadCount--;

          SuspendThread(updateThread);
     }

/*************************************************************************/
/*************************************************************************/

Notes

     I used the following definition file to create GRACE.C:

     description "Graceful Init/Deinit Example"
     output grace
     debug
     screenname "none"
     input grace z:\apps\wc386\imp31\prelude
     map
     import @z:\apps\wc386\imp31\clib.imp

     And, I used the following batch file:

     @echo off
     set inc386=z:\apps\wc386\inc31
     set wcg386=z:\apps\wc386\bin\386wcgl.exe
     @echo on

     wcc386 /3s grace
     nlmlink grace.def

     Address questions and comments to Matt Hagen, Novell, Inc., 122 East 1700
     South, Provo, Utah, 84606 (FAX 801 429-5511).

