

                          Safbench v2.31
                       Compiled 15 Jan 1998
         32bit protected mode computer, main and video memory
          speed test program for 80386+ CPUs and 387+ FPUs
                   with nice OverClocking test

              (C) Copyright 1997-1998 by Sami Farin
			All Rights Reserved
		      mailto:sfarin@ratol.fi
                     mailto:sfarin@hotmail.com
		   http://www.ratol.fi/~sfarin/





    Index
      1.   What's this?
      2.   Requirements
      3.   Usage
      4.   About Safbench
      5.   Command line arguments
      6.   Technical (?) stuff
      7.   Tables
      8.   OCTest
      9.   VESA info abbrevs
     10.   Errorlevels
     11.   Final words





[1.] - What's this?

      Safbench v2.31 is a computer speed/operation test program. It runs in
   32bit protected mode (using DOS32 v3.3 by Adam Seychell). Safbench is
   programmed completely in assembly, because FOR EXAMPLE mixing integers
   with FPU code like it's done in Safbench isn't possible in other languages
   (Simple Arithmetic FPU and Mixed FPU&CPU contain the same FPU instructions,
   but Mixed one has some integer between FPU operations).
   CPU, FPU, main memory and video memory speeds can be tested.
   For overclockers there's a test --- OCTest. It's invoked by 'Safbench O'.
   Safbench hooks the timer interrupt to achieve stabile results with
   every BIOS and (D)OS configuration, TSRs don't get called in Safbench.
   Timing accuracy is about one microsecond (assuming you don't have a buggy
   or shitty 8253/8254 timer chip). Under multitasking environments timings
   are not too precise and for example HLT-test doesn't work correctly
   ('Safbench O', then key 'H').

     Safbench is Hotmailware, it's just like Freeware, but you must send a
   message to "register_qssb@hotmail.com". That address is like a counter
   or something :) If you have something to comment etc, use sfarin@ratol.fi.
   If you don't have an email-address and use QSSB, it's ok, then you don't
   need to send an email and you can legally use QSSB :)



[2.] - Requirements

      Safbench needs 80386 or a better CPU with 387 or a better FPU.
   287 doesn't have the needed instructions needed for some tests.
   About 4MB of free mem is needed, if you use option 'M' to test
   memory speed. If you use option 'V' or 'F', Safbench needs as much memory
   as the largest video mode (for example, 1600x1200x16M needs 7500 kB)
   consumes video memory. OverClocking test needs 8224 kB.
   VBE v2.0+ is needed for 'V' and 'F'.



[3.] - Usage

      This is a VERY user-friendly program. You only have to type 'safbench'
   and press enter (assuming you are using a command interpreter).
   See chapter [5.] for more info or type 'Safbench ?'.



[4.] - About Safbench

      Safbench tests your CPU and FPU very thoroughly. It uses almost every
   instruction you've got and them are balanced so that most used instuctions
   used in the applications are also most-used in Safbench.
      But every application uses CPU and memory differently, so there can't
   be a single computer speed test program which tells the absolute and
   correct speed of your computer. Safbench tries to do it by testing
   CPU and FPU with different kind of code. It also tests memory with
   different block sizes. It reads, writes and moves in sequential
   [reverse], butterfly and random order using block sizes 1 kB - 4 MB.
   So you can test how fast your new SDRAM (or something) is.
   About the memory move routine: when it says 4 MB, it means that it
   uses about 4 MB of memory. It copies 2 MB from the start of the buffer
   to the middle of the buffer. So it reads 2 MB and writes 2 MB.
      Butterfly accessing (when reading and writing) means, that first it
   reads 32bits from start of the memory block and then from end of the
   block, then from start+4bytes and end-4bytes, start+8bytes and end-8bytes
   and so on, try to understand it... If you see butterfly read speeds which
   are faster than when using sequential reads, it's because your CPU fills
   the whole cache line when the requested data item is not present in the
   CPU's internal cache (the line is 32B on Pentiums). Butterfly routine:
    1) read from start of buffer
    2) write to end of buffer
    3) write to start of buffer+4bytes
    4) read from end of buffer-4bytes
    5) increase start buffer by 8 and decrease end buffer by 8
    6) goto 1
      Random accessing first makes a table, which contains randomly
   generated pointers to different places to memory. 'Mixed' test does
   complex memory accessing; it reads, writes and performs "read/modify"
   and "read/modify/write" instuctions (for example additions, shifts,
   and bit operations). For example, when testing 'Mixed' 4MB, Safbench
   reads one 32bit pointer (random) from memory and then reads/writes/etc
   to that position (and around it) and then gets the next pointer...
   The pointers are STORED sequentially but they point to random places
   in memory. 4MB test actually uses 8MB. MB/s value for "Random accessing"
   is calculated based on the ACTUAL numbers of BYTES accessed, including
   reading the pointers. MB/s value for Reads and Writes is based on
   bytes read or written, pointers not included. For Reads and Writes,
   the whole cache line (32 bytes) is accessed, so write speeds to memory
   are fair on CPUs which allocate new cache line on write miss (such as
   M1 and PPro/PII).
      If you want to see how fast main memory you've got (for example,
   after upgrading from FPM to SDRAM), disable L2 cache and run
   'Qsort32 16M' and 'Safbench R'.

      If you've been using CacheChk or something to test memory read speeds,
   it gives lower results, because it uses string instructions to read data
   (REP LODSD). Safbench uses a faster method. Your CPU might have L1 level
   data cache access speed less than 1 cycle per read or write, so it can do
   (as much as) two reads in parallel. Actually, Pentium+ class machines are
   capable of doing so. REP LODSD is stupid way to check cache speed, since
   NO PROGRAM uses REP LODSD to read from memory (if it does, it's a braindead
   one anyway). LODSD takes at least one cycle with every CPU and it isn't
   pairable, so you CAN'T read two double words in one cycle.
      Pentium Pro has 256 or 512 kB L2 INTERNAL cache, which is damn fast.
   A lot more faster than any conventional PB-cache sitting on a motherboard.
   It's because it runs at core speed. Pentium II's L2 cache is burst pipelined synchronous static RAM
   (BSRAM) with fast TagRAM. Transfer rates between the Pentium II processor
   core and the L2 cache are one-half the processor core clock frequency
   and scale with the processor core frequency.

      "Simple integers" tests simple CPU instructions, such as additions,
   logical operations, jumps, memory and bit manipulation. Most of them
   execute in one cycle and in parallel on a Pentium-class machine.
      "Complex integers" tests multiplications, divisions, bit scans, string
   instructions, complex memory addressing and manipulation etc. Try running
   6X86OPT if you get poor values on a Cyrix 6x86 in this test, it enables
   various options, such as N_LOCK.
      "Simple arithmetic FPU" tests additions, subtractions, multiplying,
   dividing, comparing, sign changes, loads and stores with different memory
   sizes and values. This kind of instructions are the most-used in most of
   the applications and games (such as Quake :) ).
      "Mixed FPU&CPU" tests the same instructions as the previous test, but
   does some stuffs with the integer unit while the FPU is processing data.
   386 can't do nothing but wait for the 287/387 to complete it's task, 486+
   can use the CPU while the FPU is processing data.
      "Transcendental FPU" tests transcendental instructions. Them include
   sine, cosine, tangent, arctangent, sine and cosine, 2^x-1, Y*log2(X) and
   Y*log2(X+1). 24 different values are tested for every instruction. These
   are the most time-consuming instructions you've got on your FPU.
      "Non-transcendental FPU" tests the other arithmetic instuctions not
   covered in "Simple arithmetic FPU", "Transcendental FPU" and "Processor
   control CPU" tests. For example, square root, scale, extract, remainder
   and integer part. 16 different valid values are tested with this test.
      "Processor control FPU" contains some system-managing stuff, such as
   FPU initializing, FPU state store and restore, environment saving and
   restoring, control word store/load and status word store and exception
   clearing and constant loading (0,1,pi,different log-values), BCD loads
   and stores, examine. Speed of these instructions isn't a big deal when
   comparing FPU speeds. These ARE used, not very much, though. For example
   FPU state store/restore is used in multitasking environments to store FPU
   state when switching tasks.
      "16bit speed" encodes binary data to MIME64 and does a lot of stuffs
   with mixed 8, 16 and 32 bit registers. This test shows how well your CPU
   can execute 16bit code with prefixes and such things. PPro has serious
   trouble with this test, as does PII. "Partial Register Stalls" get
   generated on them quite a lot.
      "Pentium-opt. FPU speed" tests Pentium-optimized FPU speed, it has
   90 adds, 90 subs and 30 muls. Almost half of the instructions are 'FXCH'.
   This is 'worst-case' code stream, which is not likely to be used in
   'real apps'. 'Real apps' also include main mem access and integer calc.
      "i486-opt. FPU speed" tests Intel 486 -optimized FPU speed, it does
   the SAME THINGS as Pentium-opt. code. There are no 'FXCH'-instructions.
   i486-FPU code runs a lot better on 287, 387, 486, 6x86, 6x86MX, K5,
   and K6 FPUs (because they've got the same cycle counts for latency and
   throughput). Pentium and Pentium Pro/II can achieve better results
   by scheduling FPU code (inserting 'FXCH' stuffs everywhere). C-compilers
   aren't good at optimizing/scheduling FPU code because of x86 FPU
   architecture is hard for compiler writers to model for optimization.
   Many compilers tend to use simple FPU stack model which limits their
   performance by latency rather than throughput. Thus many games
   and 3D programs, for example, are hand-optimized for Pentium FPU
   to achieve 1/2 clock throughput. That's a sad thing for other FPU's.
      "Average speed" is average of the eight tests compared to a selected
   microprocessor family (386, 486, Pentium, Cyrix 6x86, PPro, PII, K6).

      64bit precision is used in calculating FPU-speed (but it affects only
   to instructions FADD, FSUB, FMUL, FDIV and FSQRT). On Pentiums, only FDIV
   executes quicker when lower precision is used. With Cyrix 6x86, every
   FPU instruction takes the same amount of clock cycles, no matter what
   the precision is.

      Video memory speed test tests your video card memory's write speed
   and moves from main memory (or cache) to video memory. Only graphics
   linear framebuffer modes are tested. If your card doesn't support them,
   that's your problem. Without LFB, bank switching must be performed after
   every 64 kB. Using S3VBE20 or Univbe is a wise idea to increase performance,
   since bank switching is not very fast. Write speed test tests the absolute
   maximum write speed of your video card in every possible mode. The screen
   might look complex and colorful, but written data is changed only after
   the whole screen is filled. If you want to see, for example, what video
   card is "the best for "DOS games", test the cards and buy the one which
   has the best values for "Write" at the resolution and color depth you are
   using the game. Safbench doesn't test 3D speed nor GUI-acceleration speeds.
   For example (again...), Quake calculates the screen in main memory(+caches)
   and then moves it to video memory. It is where the write speed of the video
   card is needed.
      If you don't like the CLICKs when swithing video modes, turn off your
   monitor...



[5.] - Command line arguments

   Here's the possible options...

   ?: Help help help help help...

   M: Test memory speed (needs 4112 kB free mem, but if 8224 kB is available,
      also Random access speed is tested)
      See Chapter [4.] for more info (line 68).

   R: Test only Random access speed (needs 8224 kB)

   !: Use 31 different block sizes for Random accessing test when using
      'M' or 'R'

   O: Overclocking test; reads, writes and modifies memory and registers,
      tests CPU, FPU and data validity. Uses 40 different block sizes at
      randomly alternating memory positions. Writes log file after every
      five minutes and after quitted by ESC.
      Use option 'O:xxx' to specify how many seconds to test each block size,
      'O:120' tests two minutes. You can't alter log's autosave time.
      'O:M' starts from maximum memory block size and tests it forever,
      but you can still skip to next block size by pressing SPACE or 1...0
      to skip 1-10 block sizes. Use extra option ',xxx' after 'O:xxx' or
      'Q:xxx' to specify how many seconds to run OCTest (for example
      'O:60,1200').

   Q: Doesn't do parts 7-9 in OCTest (Quicksorting stuffs). With 'Q' CPU
      heats up a bit more than with 'O', since FPU is idle while Quicksorting.

   L: Blinks keyboard leds while in OCTest, only Caps Lock blinks on and off
      if everything is OK. If errors happen, also the other two start
      changing states. Leds are updated only once a second, because updating
      it so slow (at max 150 toggles/s).

   B: Test BUS speed (MHz), multiplier, MHz-rate and main mem read/write speed
      Use this on Pentiums (with or without MMX) and Cyrix 6x86MX. Sorry, but
      Bus-value isn't correct, even though it counts the number of cycles
      during which the processor's external memory bus is in use (according
      to Intel). Most of the P54C Pentiums have a bug (double issuance of
      read cycles) which will affect the results.
      Could you please tell me how to make it calculate the REAL bus speed!
      4MB block size is used for reading.
      Also number of bus cycles and CPU cycles to read/write one quadword
      (64 bits) is displayed (RAM bank is 64 bits wide).
       -bus cycles/quadword=(bus speed (MHz) / (MB/s*1,048576/8))
       -cycles/quadword=(CPU speed (MHz) / (MB/s*1,048576/8))

      Use 'B' or 'P' with option 'C' to test cache accessing (256 kB block
      size). I assume nobody has 128 kB (or less) L2 cache on a Pentium!

   P: Same as 'B', but for Pentium Pros and Pentium II's.

   V: Test video memory speed (needs VBE v2.0+ with at least one LFB mode)

   F: Same as 'V', but use 64bit accessing with FPU (intended for Pentiums)

   F:xxxx or V:xxxx tests VESA mode (in hex), for example 'F:105'

   I: Display info about every VESA mode

   S: Show every graphics linear framebuffer mode which can be tested

   W: Wait one second after mode set (default is no wait)

   0: Compare your CPU and FPU speeds to Intel 386SX 20 MHz with Cyrix FasMath

   1: Intel 486DX 33 MHz

   2: Intel Pentium 90 MHz

   3: Cyrix 6x86-P120+ 100 MHz

   4: Intel Pentium Pro 180 MHz

   5: Intel Pentium II 266 MHz

   6: AMD K6 200 MHz

      If you've got Pentium MMX or Cyrix 6x86MX, I'd be pleased to receive
   the test results! Run batch file 'MAKERES.BAT'and send it to me so I can
   include them in the next release of Safbench. With no stupid windows in the
   background. ('MAKERES1.BAT' is for P54C/P55C and Cyrix 6x86MX, 'MAKERES2.BAT'
   is for PPro and PII CPUs, 'MAKERES.BAT' is for 486, 6x86, K5, K6 etc.)
   Also K6 results for i486/P5 FPU speeds needed.
     Run the "usual" speedup programs when getting the results: 6X86OPT (ALSO
   with option '-l'), S3VBE20, MCLK, FASTVID and so on.



[6.] - Technical (?) stuff

      All the eight tests in Safbench use <4 kB of memory each, so speed
   of L2 cache nor main memory affects the result. To see how fast
   memory subsystem you've got, try command line option 'M' or run Qsort32.

      Video copy test in Safbench moves data from main memory or cache to
   video memory. Depending on how much the video mode uses memory, READ speed
   of your L1 cache, possible L2 cache and main memory limits the copy speed
   result. If you have 256 kB L2 cache, you probably get constant copy speeds
   for video modes which use less than 256 kB memory. When you test
   640x480x256c, it needs 300 kB memory, so main memory speed limits the
   maximum moving speed and speed of L2 cache doesn't matter in that case
   (very much). On my S3, 16M modes are a lot slower than other modes. Read
   speeds from video memory are not tested, because you DON'T NEED TO KNOW IT.
   If some program reads from your video card, it's a stupid program or it
   reads just a little bit, for example mouse cursor updating.

     If your processor can write to video mem in 64bit chunks (PPro with write
   combining enabled, Cyrix 6x86 with write gathering enabled or Pentium when
   FPU is used to write 64bits per one instruction), you get better results for
   both write and copy speeds in Safbench. If you get a lot better values with
   option 'F' than 'V', your processor doesn't do write gathering or combining
   OR the bandwidth to PCI write buffers (which are located between the
   processor and the PCI bus) is higher when 64bit stores are used.
   FLD [QWORD ESI+i] [1] and FSTP [QWORD EDI+i] are used for copying from main
   memory or cache to video mem and FST [QWORD EDI+i] [2] is used for writing.
   [1]: ST(0) contains a valid non-zero value
   [2]: ST(0) contains a valid non-zero value and it's increased by a big
        number after every screen written
   Video write test in Safbench is supposed to tell you the maximum throughput
   of a particular video card with every processor.
   FILD and FISTP aren't used to copy/write data, since they are a lot
   slower than FLD and FSTP. (FLD and FSTP can't be used in the "real world"
   to copy data, since a loaded value can be NaN, Denormal etc.)

   To get max speed out of your video card, I suggest you the following:
    -First, use S3VBE20 or UniVbe if VBE v2.0 isn't already supported
    -If you've got Cyrix 6x86, run 6X86OPT -linbuf. It modifies Address Region
     Register (enables write gathering for linear frame buffer). It can almost
     double the write speed to the LFB. VBE v2.0+ needed.
    -Try MCLK, overclocking might give speedups. I overclocked from 60 to 85
     MHz and it speeded up writes from 25 to 120%! Don't overclock too much,
     since it might cause lockups in win etc, when the driver doesn't get
     the data it expects to get. Speeds over 85 MHz don't give any benefits
     in anything. I noticed that using option /1 3 (use FPM timing) was faster
     than the default 2-cycle EDO (with my video card!).
     You don't have to use MCLK before you run Linux (XFree86), since there's
     a config util in which you can select all the MHz's and Hz's and stuffs :)
    -Use FASTVID or something similar proggy on a PPro if your BIOS doesn't
     turn on some bits which affect performance (DON'T use it if you have a
     buggy chipset, such as Orion 450GX rev. A2).
    -If using K6, enable write allocate with some proggy or update your BIOS.

      I use MCLK and 6X86OPT and they speed up writes to video mem from 1.7
   to 2.6 times (average being 1.95). Now my S3 Trio64 PCI runs at 85 MHz and
   writes 67 MB/s in 320x200x256c mode. That's 17 MB more than my main memory.
   Also Windows speed is increased (due to the overclocking) by 10-85%.

      About the dump 'V' and 'F' display:
   Mode=VESA mode (hexadecimal). BPP means Bits Per Pixel, 4=16 colors,
   8=256 colors, 15=32K colors, 16=64K colors, 24=16M colors, 32=16M too.
   BPSL means Bytes Per Scan Line, which when multiplied by Height is the
   number of bytes that the video mode takes memory. FPS is Frames Per Second
   (MB/s rate divided by BPSL*Height). Write and Copy speeds are in megabytes
   (1048576 bytes). Hz means vertical refresh rate in the case you didn't
   guess it! Windows(tm) Screws Up(tm) that result, too.

   The test goes like this:
   1) Switch to the video mode to be tested
   2) Fill video memory with some data
   3) Goto 2 if one second is not elapsed yet (waits for the screen to "settle
      down") if option 'W' is used
   4) Write to the video memory (clears the whole screen with different colors)
      This test lasts for about 0.3 seconds. That's enough. Testing one minute
      doesn't give any more better results.
   5) Move from main memory (or cache, depending on the mode&cache size) to
      video memory. Data is moved once before starting the timer to flush
      caches etc. This takes 0.3 seconds, too.
   6) Check Hz-rate (0.1 seconds?)
   7) Store results (0.01 microseconds?)
   8) Goto 1

   MCLK homepage: http://www.oac.uci.edu/~rliao
   S3VBE20 homepage: http://www.uni-muenster.de/math/u/mesched
   UniVbe (Display Doctor) homepage: http://www.scitechsoft.com/down_sdd.html
   TweakBios homepage: http://www.miro.pair.com/tweakbios/download2.html

      You probably get better values when the computer is "cold", just turned
   on. When it warms up, write speed of every video mode decreases by ~0-6%.
   If someone has a nice technical explanation for this little thing, I'd
   be pleased to hear it.


      FPU tests in Safbench calculate _RAW_ power of your FPU, "real world apps"
   ALWAYS include some integer and mem accessing. So the FPU-results are
   almost-worst-case-results. Pentium FPU and i486 FPU tests _ARE_ worst case
   results.

      Let's take an example. PovRay (the version 3 _I_ have) uses Pentium-
   optimized code. It means that is uses FXCH-instruction to swap the
   FPU registers so that the FPU can overlap instructions better, since
   most of the FPU-instructions need ST(0) (TOS, Top Of Stack).
   8087+ FPUs have eight registers, ST(0) to ST(7). If an instruction
   needs the result that the previous instruction used, overlapping is
   not possible. FXCH is optimized so well on Pentiums, that it can be
   executed even while the register is in use. So Pentium-optimized FPU-
   programs use FXCH quite a lot. That's a big speed decrease for other
   than Pentiums. For example, 387 needs 18 cycles for FXCH. Addition takes
   23-37 cycles. 486 needs 4 for FXCH and 8-20 for addition, Cyrix 6x86 needs
   3 and 4-9. Pentium CAN execute them BOTH or just the ADDITION in 3 cycles.
   Thus it's a big speed decrease to use Pentium-optimized code on other
   than Pentiums. Pentiums are also able to overlap FPU instructions, so
   it's possible to execute FADD, FSUB, or FMUL in one cycle (assuming the
   result for those operations is not needed in three cycles).
   Safbench uses FXCH only just a little. For the curious: Cyrix 6x86-P150+
   (2x60 MHz) runs Povray 10% faster than Pentium 75 (1.5x50 MHz).
      So... PovRay would run faster on 387, 486 and Cyrix 6x86 if it didn't
   use FXCH to achieve speedups on Pentiums (yes, I know there are still
   compilers which support 486-optimizations).

      Oh well. Then there's Quake. That's also an example of a program, which
   uses FPU very extensively, and it's optimized for Pentiums. So you are
   probably not surprised that it performs poorly on Cyrix 6x86. Please
   note that Pentium 120 runs at 120 MHz, but Cyrix 6x86-P120+ runs at 100 MHz.
   If Quake was optimized for, let's say, 486, Cyrix 6x86-P120+ wouldn't seem
   so slow when compred to Pentium 100. Quake also does integer calculation.
   Quake is so fast on Pentium Pros, because it's got VERY fast L2 cache (read
   speed about 400 MB/s on PPro/180) and CPU and FPU are at least 20% faster
   than on Pentiums (with same MHz-rate). PPro has a nice technique, which
   allows 64bit moves in one cycle with REP MOVSD instruction.
      I've read articles about Quake and Cyrix 6x86, where was stated, that
   Quake uses FPU only to BUFFER the data, but I don't agree. After debugging
   Quake, I found that it actually did something 'useful', such as multiplying
   and addition. I didn't see it move data via the FPU in 64bit chunks.
   Actually, it doesn't move data all day long, except to video mem.
   Cyrix 6x86 does write gathering (so data is moved in 64bit chunks anyway,
   even though a software uses 32 or 16bit move), so it's best to use CPU to
   move the data (REP MOVSD)... And the same goes for Pentium Pros.
   But Pentium doesn't do write gathering, so it might be faster to use
   the FPU to move data. Quake is optimized purely for Pentiums, it takes
   advantage of FXCH being pairable with some FPU instructions, so it can
   be executed with no extra cost. On a 6x86 FXCH takes 3 clocks.
      Also, main routine in Quake is Pentium-optimized assembly-code, for
   example FDIV with lowest precision. That FDIV is parallelized with integer
   ops (C-compilers can't make that kind of optimized code). So, when FDIV
   takes 39 cycles on a Pentium with highest precision (the default one, 64
   bits), it takes 34 cycles on Cyrix 6x86. But in Quake, 24 bit precision
   is used. With Cyrix 6x86, that FDIV still takes 34 cycles, but with
   Pentium only 19 cycles. Quake is 100% FPU down to the 8 or 16 pixel sub-
   divisions. Quake is optimized so that FPU and integer units complete
   roughly at the same time, and that piece of code is timed for Pentiums!
   Quake is not a GOOD benchmark of FPU speed; it's a GOOD benchmark of
   Quake speed.
   See http://www.gamers.org/dEngine/quake/papers/mikeab-cgdc.html for more
   information.

      My Cyrix 6x86-P150+ at 120 MHz runs PovRay 10% faster than Pentium at
   75 MHz, so FPU-speed of Cyrix isn't very good. In contrast, CPU is a lot
   faster. Quake isn't the only program on the world. 6x86-P150+ is just as
   fast as Pentium 75 in Quake (tested with the same mobo).

     And also Excel uses FPU (long reals, 64bits) in the calculations...
   But there's so much overhead that speed of FPU doesn't matter very much.


   Some size statistics of Safbench.EXE:
    - Main code  10.2 kB (timing routines, number displaying etc)
    - Mem speeds 10.2 kB (main and video memory speed tests)
    - Test code   7.7 kB (Simple Int., Complex Int., FPU stuffs etc.)
    - Data       11.5 kB (FPU test data (long reals etc), texts etc.)
    - OCTest      7.9 kB (Only code size.)
   The executable is compressed and it contains the stub loader.



[7.] - Tables

   Here's the results I've got from different computers
   (Cyrix 6x86MX-results still missing!).
   (P54=Pentium, P55=Pentium MMX, CxM1/100=Cyrix 6x86-P120+)

Nr Test
1: Simple integers
2: Complex integers
3: Simple arithmetic FPU
4: Mixed FPU&CPU
5: Transcendental FPU
6: Non-transcendental FPU
7: Processor control FPU
8: 16bit speed
p: Pentium-optimized FPU
i: Intel 486-optimized FPU

     386/20   486/33    P54/90   P55/200  CxM1/100  PPro/180   PII/266
1:   6727.3  40723.9  181672.9  334841.6  220270.9  410091.1  610191.8
2:   4728.3  18169.5   61236.7  125960.7   84934.8  166391.7  246439.9
3:   2815.7  12390.1  108807.4  240844.8   69841.8  259873.3  385103.6
4:   2132.7  11258.3   83058.6  183732.5   66460.9  183286.0  272998.9
5:    446.4    602.2    4780.2   10593.6    3395.4    9738.1   14408.6
6:   2997.2   7018.6   26390.0   58381.9   32720.5   56537.9   84124.6
7:   4877.3  17695.8   60254.6  130738.3   70579.9   87406.1  134912.8
8:   2609.2  10525.4   46339.9  100286.4   56394.5   53254.8   79843.7
p:   1346.8   9237.4  260697.6  593178.0   39913.4  500334.9  744471.6
i:   2004.8  12067.9  116089.3  260627.0   69454.5  414824.6  617235.9

     K6/200
   409915.5
   142768.1
   207291.2
   158958.5
     9253.9
    72311.5
   101341.5
    86928.4
        ?.? NEEDED
        ?.? NEEDED

   K6 and 6x86MX owners might want to give me some results...?


Power Per MHz
        386      486      P54      P55  Cyrix M1      PPro      PII   AMD K6
1:    336.4   1221.7   2018.6   1674.2    2222.2    2262.8   2276.5*  2049.6
2:    236.4    545.1    680.4    629.8     837.3     902.5    919.4*   713.8
3:    140.8    371.7   1209.0   1204.2     698.4    1419.5   1436.8*  1036.5
4:    106.6    337.8    922.9    918.7     664.6    1013.1   1018.5*   794.8
5:     22.3     18.1     53.1     53.0      34.0      54.1*    53.8     46.3
6:    149.9    210.6    293.2    291.9     327.2     313.2    313.9    361.6*
7:    243.9    530.9    669.5    653.7     705.8*    483.7    503.4    506.7
8:    130.4    315.8    514.9    501.4     563.9*    295.1    297.9    434.6
p:     67.3    277.1   2896.6   2965.9*    399.1    2779.6   2779.6      ?.? ???
i:    100.2    362.0   1289.9   1303.1     694.5    2304.6*  2304.6*     ?.? ???
*=most powerful

   Intel Processor families, processor speed improvements
                  From...  386      486     Pentium    PPro
                  To.....  486    Pentium     PPro      PII
Simple integers           3.87x    1.65x     1.13x    1.01x
Complex integers          2.39x    1.25x     1.39x    1.02x
Simple arithmetic FPU     2.66x    3.25x     1.18x    1.01x
Mixed FPU&CPU             3.18x    2.73x     1.10x    1.01x
Transcendental FPU        0.80x    2.93x     1.02x    1.00x
Non-transcendental FPU    1.44x    1.39x     1.07x    1.00x
Processor control FPU     2.19x    1.26x     0.74x    1.04x
16bit speed               2.40x    1.63x     0.59x    1.01x
Pentium FPU               4.12x   10.45x     0.96x    1.00x
i486 FPU                  3.61x    3.56x     1.79x    1.00x
-------------------------------------------------------------
Average                   2.67x    3.01x     1.10x    1.01x


   As you can see, difference between PPro and PII is very little,
   but PII has 32 kB L1 cache (PPro has 16 kB) and IA MMX.
   PII's L1 data cache is 4-way set-associative (16 kB), PPro's L1
   data cache is is 2-way set-associative (8 kB).
   PII's L2 cache runs at half the core speed (133 MHz for 266
   MHz CPU), PPro's L2 cache runs as fast as the CPU (but it
   can't be accessed as fast as L1 cache, see below).
   PII's maximum core frequency is higher than that of PPro's.


   Memory speed results (MB/s), speeds taken from cache boundary
   (8kB and 64kB for 486/33 etc.) If the result is in brackets, it's
   not necessarily correct and I want results for that particular CPU!

   M1/100a: Cyrix 6x86-P120+ 100 MHz, MG mobo
   M1/100b: Cyrix 6x86-P120+ 100 MHz, Asus VX97, 16 MB FPM
   M1/120: Cyrix 6x86-P150+ 120 MHz, Asus VX97, 16 MB FPM
   M1/120e: Cyrix 6x86-P150+ 120 MHz, Asus VX97, 64 MB EDO
   M1/133: Cyrix 6x86-P166+ 133 MHz   (no L2 cache)
   M1/133t: Cyrix 6x86-P166+ 133 MHz  (ASUS mobo, 512 kB PB-cache, 64 MB EDO)
   M1/150:  Cyrix 6x86-P200+ 150 MHz  (ASUS mobo, 512 kB PB-cache, 64 MB EDO)
   AMD K6/200: 66.666*3
   AMD K6/208: 83.333*2.5=208,333
   AMD K6/208t: same as above, but memory timings tweaked
   PII/266: Osborne Pentium II 266 MHz, 2*32 MB BEDO, 512 kB cache

                          Sequential accessing
                   read                write               move
   CPU         L1    L2  main      L1    L2  main      L1    L2  main
   486/33    70.8  35.9  12.9    61.4  62.4  30.7    41.5  29.5   7.9
  [P54/90   313.0 107.6  64.0   448.8  47.2  28.8   246.3  37.4  18.7]
   P54/120  718.2 178.1  89.3   389.7  78.2  73.3   396.5  61.0  34.9
   P54/166 1002.2 202.3 100.5   473.7  86.7  83.5   540.4  67.6  42.7
   P55/200 1243.7 223.6 152.9  1343.6! 86.9  84.7   700.4  69.7  48.1
   P55/225 1402.7 251.2 171.8  1509.2! 97.6  95.1   789.5  78.3  54.0
   P55/233 1439.3 251.0 152.9  1567.6! 86.9  84.7   805.7  69.7  48.1
   M1/100a  607.3 194.1  72.2   288.8 109.2  50.4   376.8  66.1  26.4
   M1/100b  607.3 192.6  72.2   288.8 113.2  50.4   376.8  72.8  26.7
   M1/120   712.1 230.7  86.7   345.9 135.6  60.5   450.7  87.2  32.0
   M1/120e  712.1 230.7 107.0   345.9 135.6  67.2   450.7  87.2  36.4
   M1/133  *638.6 -----  74.7   376.3 -----  41.2   494.4 -----  26.6
   M1/133t  815.1 256.2 111.7   384.0 150.6  64.8   500.4  96.8  37.5
   M1/150   896.4 287.7 125.4   429.8 169.1  72.8   562.0 108.7  42.2
   K6/200   731.6 250.1 125.4   725.9 127.0  69.3   748.0  84.3  42.3
   K6/208   763.0 282.2 125.3   758.1 154.0  73.8   781.2 101.1  44.2
   K6/208t  763.0 282.2 147.8   758.1 154.0  86.8   781.2 101.1  51.0
   P6/180   672.0 394.8 196.4   545.7 342.8  71.9   759.0 222.0  43.9
   PII/266 1001.1 479.5 222.3   811.1 239.0  73.9  1233.7 182.9  51.6

   *: 847.6 with block size of 8 kB
   !: With block size of 8 kB

                          Reverse accessing
                   read                write               move
   CPU         L1    L2  main      L1    L2  main      L1    L2  main
   486/33    71.0  37.3  13.1    55.4  62.5  30.7    40.6  29.5   8.1
  [P54/90   313.0 107.7  65.7   447.9  47.2  28.9   278.9  37.4  19.1]
   P54/120  716.9 178.1  89.3   382.2  78.2  74.9   406.2  61.0  34.9
   P54/166  998.2 202.3 100.5   471.9  86.7  83.5   546.9  67.6  42.7
   M1/100a  624.2 123.0  53.7   285.0 109.1  50.4   284.2  54.6  23.9
   M1/100b  624.2 133.2  53.7   285.0 109.2  50.4   284.2  61.2  24.3
   M1/120   726.2 159.6  64.7   341.6 130.8  60.5   340.4  73.2  29.1
   M1/133   719.9 -----  57.3   371.5 -----  41.2   382.9 -----  24.0
   M1/133t  819.8 177.2  79.3   379.2 145.2  64.8   377.9  81.3  34.8
   M1/150   897.0 199.4  89.1   424.4 163.0  72.8   424.6  91.3  39.1
   K6/200   746.3 254.0 125.4   732.0 127.0  69.3   747.7  84.3  42.3
   K6/208   778.0 281.5 125.4   763.7 154.0  73.8   781.2 101.1  44.2
   K6/208t  778.0 282.2 147.9   763.7 154.0  86.8   781.2 101.1  51.0
   P6/180   674.0 390.8 182.8   545.7 351.2  71.9   293.0 221.1  36.3


                          Butterfly accessing
                   read                write               move
   CPU         L1    L2  main      L1    L2  main      L1    L2  main
   486/33    70.4  40.8  10.5    60.7  60.5  15.4    34.7  21.2   6.2
  [P54/90   299.0 110.8  53.3   365.3  47.0  20.3   194.1  38.1  16.6]
   P54/120  415.7 157.1  82.0   316.0  78.1  23.7   310.4  57.1  25.7
   P54/166  578.3 192.4  98.0   384.0  86.6  29.2   430.7  63.3  30.3
   P55/200  721.7 263.3 115.1  1159.3! 86.8  34.2   534.9  67.3  34.5
   P55/225  800.0 295.7 129.4  1302.1! 97.4  38.5   602.0  75.6  38.8
   P55/233  826.5 267.7 118.3  1352.6! 86.8  34.2   617.4  67.3  34.5
   M1/100a  670.6 151.0  59.6   190.0  92.6  47.0   186.2  47.8  22.8
   M1/100b  670.6 167.4  60.9   190.0  95.5  47.2   186.2  56.6  23.6
   M1/120   804.7 200.5  72.9   227.7 114.4  56.6   223.8  67.8  28.3
   M1/120e  804.7 200.5  88.7   227.7 114.4  62.5   223.8  67.8  31.2
   M1/133   833.3 -----  69.1   249.1 -----  41.2   238.7 -----  19.0
   M1/133t  894.4 222.6 100.5   252.7 127.0  64.5   248.6  75.3  32.2
   M1/150  1035.5 250.1 112.9   283.2 142.7  72.5   277.5  84.5  36.2
   K6/200   734.8 233.8 124.8   726.6 126.0  68.7   577.7  66.3  34.4
   K6/208   766.2 279.6 129.6   758.7 152.8  73.2   604.6  80.3  36.7
   K6/208t  766.2 279.6 153.9   758.7 152.8  85.8   604.6  80.3  43.0
   P6/180   674.0 475.9 113.1   494.3 326.3  53.3   208.7 112.2  29.0
   PII/266  998.7 540.6 134.6   732.4 238.7  63.1   506.2 119.4  33.5

   !: With block size of 8 kB


   These are for Safbench v1.32:
                         Random accessing
   Cyrix 6x86-P120+ (256 kB PB, 16 MB FPM RAM, MG i430VX mobo)
   With SADS enabled/disabled, speedup displayed for Mixed:
   Block size    Mixed                     Read             Write
   1        195.428/195.425          423.692/423.692   263.923/263.923
   2        197.336/197.339          425.273/425.273   261.743/261.743
   4        199.942/199.923          425.411/425.411   261.158/261.162
   8        188.310/188.359          343.730/343.541   260.116/260.135
   16       135.166/137.681  1.86%   191.103/191.387   146.147/147.752
   32        78.132/ 84.335  7.94%   146.816/146.911   102.537/104.844
   64        61.998/ 68.398 10.32%   131.345/131.507    89.535/ 91.973
   128       55.776/ 62.020 11.19%   124.991/125.132    84.321/ 86.866
   256       52.862/ 58.861 11.35%   120.015/120.203    81.484/ 84.133
   512       34.213/ 36.142  5.64%    84.212/ 84.256    58.376/ 59.909
   1024      27.671/ 28.684  3.66%    68.967/ 69.064    49.333/ 50.458
   2048      25.149/ 25.946  3.17%    62.399/ 62.573    45.292/ 46.308
   4096      22.576/ 23.281  3.12%    58.526/ 58.778    42.170/ 43.392
   Avg.:     98.043/100.484  2.49%   200.499/200.510   134.318/135.584


   AMD K6/83.333*2.5 MHz, ASUS P/I P55T2P4, 512 kB PB-cache, 64 MB EDO
   Block size    Mixed         Read        Write
   1           373.015      629.963      773.795
   2           377.418      630.913      773.285
   4           379.344      630.875      775.466
   8           377.372      631.245      775.632
   16          372.540      623.789      759.802
   32          237.125      327.785      278.042
   64          112.521      233.931      174.797
   128          93.360      211.812      156.690
   256          86.757      203.742      150.902
   512          83.660      198.390      147.052
   1024         57.710      130.714       99.605
   2048         49.193      111.404       85.822
   4096         43.246      103.595       80.371
   Avg.:       203.328      359.089      387.020


   Cyrix 6x86-P166+ (133 MHz), ASUS P/I P55T2P4, 512 kB PB-cache, 64 MB EDO
   Block size    Mixed         Read        Write
   1           259.858      563.387      350.938
   2           262.393      565.466      348.041
   4           265.744      565.676      347.261
   8           250.740      457.401      345.923
   16          186.419      266.773      206.649
   32          115.155      206.706      147.252
   64           92.936      185.025      128.697
   128          84.131      176.046      121.262
   256          80.032      171.752      117.875
   512          78.066      168.261      115.423
   1024         47.724      115.500       80.337
   2048         38.604       98.192       67.480
   4096         33.115       89.544       61.107
   Avg.:       138.071      279.210      187.557


   Cyrix 6x86-P200+ (150 MHz), ASUS P/I P55T2P4 mobo, 512 kB PB-cache, 64 MB EDO
   Block size    Mixed         Read        Write
   1           291.884      632.785      394.173
   2           294.726      635.164      390.932
   4           298.562      635.384      390.042
   8           281.628      512.774      388.527
   16          208.419      293.247      228.372
   32          128.588      223.527      163.267
   64          103.892      202.029      143.035
   128          94.048      194.296      134.942
   256          89.664      190.943      131.739
   512          87.718      188.968      129.843
   1024         53.730      129.747       90.740
   2048         43.418      110.273       76.088
   4096         37.219      100.591       67.861
   Avg.:       154.884      311.518      209.966



   These are for Safbench v2.13+:
                         Random accessing

   Pentium MMX/200 512 kB L2 cache (66 MHz BUS)
   Block size   Mixed        Read       Write
   1           190.26     1001.92     1001.95
   2           208.21     1002.91     1002.90
   4           192.93     1004.68     1004.68
   8           164.17      922.09      778.22
   16          117.26      493.63      154.20
   32           86.02      260.43       97.30
   64           71.87      205.03       83.81
   128          66.06      183.89       82.55
   256          62.53      172.46       81.95
   512          55.82      119.81       69.57
   1024         40.80       87.63       61.36
   2048         35.49       78.26       59.35
   4096         33.23       74.44       58.38
   Avg.:       101.90      431.32      348.94


   Pentium MMX/225 512 kB L2 cache (75 MHz BUS)
   Block size   Mixed        Read       Write
   1           212.67     1125.38     1125.42
   2           232.20     1126.49     1126.49
   4           216.40     1128.51     1128.48
   8           184.17     1021.06      770.67
   16          131.51      555.07      170.44
   32           96.37      292.57      109.37
   64           80.62      230.23       94.06
   128          74.23      206.59       92.72
   256          70.16      193.75       92.05
   512          62.65      134.57       78.14
   1024         45.84       98.44       68.92
   2048         39.87       87.93       66.66
   4096         37.33       83.65       65.57
   Avg.:       114.16      483.40      383.77


   Pentium MMX/233 512 kB L2 cache (66 MHz BUS)
   Block size   Mixed        Read       Write
   1           219.64     1168.91     1168.96
   2           241.60     1170.10     1170.04
   4           222.12     1172.12     1172.18
   8           183.89     1049.52      760.64
   16          125.59      537.26      153.01
   32           89.40      268.34       97.53
   64           73.87      208.44       83.75
   128          67.52      185.97       82.54
   256          63.72      174.19       81.95
   512          56.84      123.24       69.57
   1024         41.30       88.88       61.36
   2048         35.85       78.82       59.35
   4096         33.56       74.72       58.38
   Avg.:       111.92      484.66      386.10


   Pentium Pro/180 256 kB L2 cache
   Block size    Mixed         Read        Write
   1           156.615      544.836      442.946
   2           159.211      541.359      452.735
   4           148.089      542.610      454.725
   8           135.329      479.488      435.596
   16          126.346      408.666      403.304
   32          118.767      377.808      374.902
   64          114.498      364.155      352.010
   128         111.080      352.769      333.658
   256         100.194      194.794      184.965
   512          53.379      129.550       83.724
   1024         39.450      112.402       63.332
   2048         34.700      105.567       56.084
   4096         32.701      102.329       53.025
   Avg.:       102.335      327.410      283.923


   Pentium II/266 512 kB L2 cache
   Block size   Mixed        Read       Write
   1           329.99      810.65      659.09
   2           345.15      808.22      673.66
   4           321.58      809.22      678.26
   8           303.88      777.21      654.30
   16          255.11      620.97      545.64
   32          196.00      506.00      373.29
   64          163.12      461.16      290.34
   128         149.22      440.21      256.15
   256         141.03      424.56      241.24
   512         123.21      214.98      186.31
   1024         68.63      149.61      103.97
   2048         50.03      130.51       76.37
   4096         43.59      123.01       66.63
   Avg.:       191.58      482.79      369.64


   Cyrix 6x86-P120+ (256 kB PB, 16 MB FPM RAM, MG i430VX mobo)
   With SADS enabled/disabled, speedup displayed for Mixed and Write:
   Block size    Mixed                    Read              Write
   1        171.830/171.835         423.692/423.680    264.402/264.402
   2        164.746/165.080  0.20%  426.327/426.327    263.845/263.841
   4        153.398/155.446  1.34%  426.399/426.217    263.618/263.618
   8        134.603/137.894  2.44%  401.759/401.753    262.177/262.221
   16        97.326/102.613  5.43%  265.927/266.133    179.422/182.025 1.45%
   32        71.016/ 77.150  8.64%  165.790/166.132    112.048/114.557 2.24%
   64        59.432/ 65.381 10.01%  138.675/138.837     93.133/ 95.572 2.62%
   128       54.611/ 60.417 10.63%  128.084/128.294     85.975/ 88.511 2.95%
   256       51.437/ 56.844 10.51%  120.194/120.320     81.453/ 84.080 3.23%
   512       33.234/ 34.956  5.18%   84.216/ 84.261     58.384/ 59.922 2.63%
   1024      27.318/ 28.220  3.30%   69.113/ 69.212     48.956/ 50.145 2.43%
   2048      24.884/ 25.565  2.74%   62.672/ 62.826     44.975/ 46.046 2.38%
   4096      23.626/ 24.228  2.55%   58.548/ 58.797     42.190/ 43.413 2.90%
   Avg.:     82.112/ 85.048  3.58%  213.184/213.291    138.506/139.873 0.99%


   Cyrix 6x86-P120+ (512 kB PB, 16 MB FPM RAM, Asus VX97 mobo)
   With SADS disabled
   Block size    Mixed         Read        Write
   1           156.698      423.692      264.402
   2           164.844      425.273      263.845
   4           156.069      425.405      263.618
   8           137.380      401.184      262.881
   16          102.941      264.983      182.782
   32           77.414      165.828      114.346
   64           65.629      138.679       95.713
   128          60.524      128.265       88.488
   256          58.222      123.630       85.279
   512          56.380      119.600       83.045
   1024         34.106       78.732       58.451
   2048         27.946       66.389       49.804
   4096         25.395       60.707       45.219
   Avg.:        86.427      217.105      142.913


   Cyrix 6x86-P133+ (512 kB PB, 16 MB FPM RAM, Asus VX97 mobo)
   With SADS enabled/disabled, speedup displayed for Mixed:
   Block size    Mixed                     Read             Write
   1        186.857/186.857          465.806/465.806   289.915/289.911
   2        193.002/193.005          468.201/468.201   287.444/287.449
   4        180.912/181.519          468.073/468.280   286.806/286.793
   8        149.810/152.671  1.9%    439.982/439.989   288.285/287.877
   16       106.913/112.891  5.6%    290.747/290.787   196.126/199.032
   32        77.173/ 83.880  8.7%    182.065/182.314   123.003/125.762
   64        64.863/ 71.261  9.9%    152.389/152.518   102.370/105.099
   128       59.382/ 65.650 10.6%    140.741/140.966    94.420/ 97.234
   256       56.966/ 63.057 10.7%    135.646/135.853    90.719/ 93.628
   512       55.154/ 61.042 10.7%    131.231/131.455    88.238/ 91.227
   1024      35.283/ 37.159  5.3%     86.595/ 86.751    61.637/ 63.755
   2048      29.396/ 30.438  3.5%     73.031/ 73.235    52.880/ 54.419
   4096      26.819/ 27.607  2.9%     66.421/ 66.728    48.071/ 49.615
   Avg.:     94.041/ 97.464  3.6%    238.533/238.683   154.609/156.292


   Cyrix 6x86-P150+ (512 kB PB, 16 MB FPM RAM, Asus VX97 mobo)
   With SADS enabled/disabled:
   Block size    Mixed             Read              Write
   1        191.665/191.980   507.504/507.481   316.707/316.703
   2        204.631/204.890   509.389/509.374   316.040/316.040
   4        190.112/191.857   509.540/509.571   315.743/315.752
   8        164.311/167.947   480.305/480.312   313.994/313.994
   16       117.256/123.504   316.376/316.813   215.821/218.889
   32        85.204/ 92.590   198.203/198.430   133.813/136.887
   64        71.474/ 78.545   166.118/166.350   111.649/114.590
   128       65.644/ 72.453   153.398/153.611   102.937/106.021
   256       62.935/ 69.600   147.868/148.171    98.966/102.139
   512       60.934/ 67.375   143.007/143.222    96.303/ 99.559
   1024      38.809/ 40.876    93.855/ 94.036    67.775/ 70.009
   2048      32.359/ 33.488    79.201/ 79.454    57.931/ 59.567
   4096      29.564/ 30.423    72.426/ 72.755    52.491/ 54.163
   Avg.:    101.146/105.041   259.784/259.968   169.244/171.101


   Same as above, but with 64 MB EDO (SADS disabled), speed difference displ.:
   Block size   Mixed          Read         Write
   1           189.49        507.50        316.70
   2           205.41        509.39        316.04
   4           191.50        509.56        315.77
   8           168.47        480.05        314.91
   16          127.22        317.50        218.36
   32           96.08        209.04        141.56
   64           81.60        177.93        119.52
   128          75.53        165.53        110.94
   256          72.49  4.2%  160.03  8.0%  107.00  4.8%
   512          70.41        155.52        104.52
   1024         46.12        109.44         76.27
   2048         38.58         93.73         65.68
   4096         35.29 16.0%   85.81 17.9%   59.67 10.2%
   Avg.:       107.55  2.4%  267.77  3.0%  174.38  1.9%


  "The ADS# (Address Strobe) signal is asserted to
   indicate the validity of the transaction address on the
   A[35:3]# pins. All bus agents observe the ADS#
   activation to begin parity checking, protocol
   checking, address decode, internal snoop, or
   deferred reply ID match operations associated with
   the new transaction. This signal must connect the
   appropriate pins on all Pentium II processor System
   Bus agents."
   That's from Intel's Pentium II docs about ADS pin.


   In Mixed test, the whole cache line is often fetched from cache/main
   mem and then processed in CPU's cache.
   With Random reads/writes, 32 bytes are read/written from/to a random
   place and then proceeded to the next random position. When cache read
   miss occurs, CPU fetches the whole cache line (32B) from main memory
   on Pentium+ CPUs, but when cache write miss occurs, Pentium doesn't
   allocate new cache line for it (get it from memory), but PPro does.
   That's why 32 byte reads AND writes are used.
   The randomization routine gives the same values for every run with
   every CPU, so memory is accessed at the same places every time you
   run Random access test.

   Main memory speeds are taken from 4096 kB block size.



   Speed of my Diamond Stealth 64 DRAM PCI, 1 MB 2-cycle EDO, Trio64 (764).
   Main memory is 2*8 MB 60 ns FPM RAM, 256 kB PB-cache, MG i430VX mobo.
   Fastest possible settings for RAM in BIOS. Cyrix 6x86-P120+ (100 MHz).
   The shitty monitor is ADI MicroScan 2E, about 15 inches.

---Clear boot---
S3VBE20, no MCLK run (S3's DRAM clock generator running at 59.96 MHz):
Mode   Width  Height  BPP   BPSL   Write      FPS    Copy      FPS      Hz
 163     320     200    8    320   35.40   580.02   33.23   544.50   69.94
 164     320     240    8    320   31.79   434.08   30.71   419.25   67.27
 165     320     400    8    320   35.36   289.69   33.23   272.23   69.94
 166     320     480    8    320   31.72   216.57   30.67   209.34   67.27
 14F     400     300    8    400   37.16   324.75   33.23   290.37   73.40
 12D     512     384    8    512   36.75   195.99   33.23   177.22   81.60
 100     640     400    8    640   36.76   150.55   33.20   136.00   69.87
 101     640     480    8    640   36.74   125.39   30.84   105.26   60.16
 103     800     600    8    800   33.96    74.18   26.82    58.60   56.15
 105    1024     768    8   1024   25.25    33.66   22.84    30.45   87.10
 110     640     480   15   1280   30.18    51.50   26.32    44.92   60.04
 111     640     480   16   1280   30.17    51.48   26.32    44.92   60.04
 113     800     600   15   1600   22.71    24.80   20.04    21.89   56.15
 114     800     600   16   1600   22.72    24.82   20.04    21.89   56.15
 211     640     400   32   2560   15.98    16.36   12.24    12.53   70.09


---MCLK used---
S3VBE20, MCLK /0 93 2 2 /1 3 (85.01 MHz):
Mode   Width  Height  BPP   BPSL   Write      FPS    Copy      FPS      Hz
 163     320     200    8    320   37.30   611.12   33.23   544.50   69.94
 164     320     240    8    320   37.30   509.26   33.23   453.75   67.27
 165     320     400    8    320   37.30   305.56   33.23   272.23   69.94
 166     320     480    8    320   37.30   254.63   33.23   226.86   67.27
 14F     400     300    8    400   37.36   326.45   33.23   290.37   73.40
 12D     512     384    8    512   37.30   198.93   33.23   177.22   81.60
 100     640     400    8    640   37.30   152.78   33.20   136.00   69.87
 101     640     480    8    640   37.30   127.32   30.84   105.26   60.16
 103     800     600    8    800   37.30    81.48   26.84    58.63   56.14
 105    1024     768    8   1024   37.30    49.73   26.32    35.09   87.09
 110     640     480   15   1280   37.30    63.66   26.32    44.92   60.04
 111     640     480   16   1280   37.30    63.66   26.32    44.92   60.04
 113     800     600   15   1600   34.98    38.21   26.31    28.74   56.15
 114     800     600   16   1600   34.96    38.18   26.31    28.74   56.15
 211     640     400   32   2560   21.28    21.79   17.48    17.90   70.09


---6X86OPT used---
6X86OPT -LINBUF, S3VBE20, no MCLK run (59.96 MHz):
Mode   Width  Height  BPP   BPSL   Write      FPS    Copy      FPS      Hz
 163     320     200    8    320   43.52   713.03   40.11   657.24   69.94
 164     320     240    8    320   38.54   526.21   35.89   489.96   67.27
 165     320     400    8    320   43.27   354.43   39.93   327.08   69.94
 166     320     480    8    320   38.36   261.88   35.67   243.49   67.27
 14F     400     300    8    400   43.85   383.14   41.62   363.67   73.40
 12D     512     384    8    512   46.78   249.51   42.21   225.14   81.60
 100     640     400    8    640   44.76   183.34   41.29   169.13   69.87
 101     640     480    8    640   43.82   149.56   39.48   134.76   60.16
 103     800     600    8    800   39.88    87.12   34.24    74.81   56.15
 105    1024     768    8   1024   30.42    40.56   25.23    33.64   87.07
 110     640     480   15   1280   37.35    63.75   30.09    51.35   60.04
 111     640     480   16   1280   37.37    63.78   30.10    51.36   60.04
 113     800     600   15   1600   28.52    31.15   22.60    24.69   56.14
 114     800     600   16   1600   28.57    31.20   22.60    24.69   56.15
 211     640     400   32   2560   26.70    27.34   16.05    16.44   70.09


---6X86OPT and MCLK used---
6X86OPT -LINBUF, S3VBE20, MCLK /0 93 2 2 /1 3 (85.01 MHz):
Mode   Width  Height  BPP   BPSL   Write      FPS    Copy      FPS      Hz
 163     320     200    8    320   69.15  1132.97   52.70   863.46   69.94
 164     320     240    8    320   65.63   896.07   52.70   719.59   67.27
 165     320     400    8    320   69.03   565.46   52.71   431.77   69.94
 166     320     480    8    320   65.49   447.09   52.71   359.84   67.27
 14F     400     300    8    400   63.29   553.01   52.70   460.53   73.40
 12D     512     384    8    512   63.61   339.26   52.70   281.08   81.60
 100     640     400    8    640   62.18   254.68   52.64   215.63   69.87
 101     640     480    8    640   61.81   210.97   46.94   160.23   60.16
 103     800     600    8    800   58.09   126.89   38.27    83.61   56.15
 105    1024     768    8   1024   64.62    86.16   37.12    49.50   87.07
 110     640     480   15   1280   53.06    90.55   37.22    63.52   60.04
 111     640     480   16   1280   53.06    90.56   37.22    63.52   60.04
 113     800     600   15   1600   50.83    55.52   34.79    38.00   56.15
 114     800     600   16   1600   51.25    55.97   34.77    37.98   56.15
 211     640     400   32   2560   33.79    34.60   21.25    21.76   70.09


---6X86OPT and MCLK used, Cyrix 6x86 overclocked to 120 MHz---
----(CPU, main memory, L1, and L2 cache now 20% faster)-----
6X86OPT -LINBUF, S3VBE20, MCLK /0 93 2 2 /1 3 (85.01 MHz):
Mode   Width  Height  BPP   BPSL   Write      FPS    Copy      FPS      Hz
 163     320     200    8    320   74.29  1217.20   63.20  1035.39   69.94
 164     320     240    8    320   70.15   957.76   63.18   862.66   67.27
 165     320     400    8    320   74.07   606.77   63.19   517.68   69.94
 166     320     480    8    320   69.91   477.25   63.19   431.40   67.27
 14F     400     300    8    400   67.56   590.37   60.54   528.96   73.40
 12D     512     384    8    512   68.97   367.87   60.39   322.09   81.60
 100     640     400    8    640   67.05   274.63   59.99   245.71   69.87
 101     640     480    8    640   66.28   226.24   54.47   185.92   60.15
 103     800     600    8    800   61.33   133.97   45.52    99.44   56.14
 105    1024     768    8   1024   65.31    87.08   43.23    57.65   87.06
 110     640     480   15   1280   57.31    97.81   44.66    76.22   60.04
 111     640     480   16   1280   57.34    97.86   44.66    76.23   60.04
 113     800     600   15   1600   51.40    56.15   37.13    40.56   56.14
 114     800     600   16   1600   51.40    56.14   37.11    40.53   56.14
 211     640     400   32   2560   39.15    40.09   24.06    24.64   70.09


Same as above, but with Asus VX97 mobo (512 kB cache)
Mode   Width  Height  BPP   BPSL   Write      FPS    Copy      FPS      Hz
 163     320     200    8    320   73.67  1207.02   71.50  1171.51   69.95
 164     320     240    8    320   69.54   949.46   67.77   925.32   67.28
 165     320     400    8    320   73.41   601.35   71.21   583.36   69.95
 166     320     480    8    320   69.28   472.98   67.52   460.93   67.28
 14F     400     300    8    400   67.77   592.19   65.67   573.85   73.41
 12D     512     384    8    512   69.21   369.12   66.42   354.23   81.62
 100     640     400    8    640   67.02   274.53   64.80   265.44   69.88
 101     640     480    8    640   66.27   226.20   64.09   218.77   60.16
 103     800     600    8    800   61.62   134.62   59.72   130.46   56.16
 105    1024     768    8   1024   65.32    87.10   47.42    63.22   87.08
 110     640     480   15   1280   56.93    97.16   52.38    89.39   60.05
 111     640     480   16   1280   56.93    97.16   52.38    89.39   60.05
 113     800     600   15   1600   51.41    56.16   39.34    42.97   56.15
 114     800     600   16   1600   51.41    56.16   39.32    42.95   56.15
 211     640     400   32   2560   39.10    40.03   26.89    27.54   70.10


Same as above, but with Asus VX97 mobo (512 kB cache), MCLK /4 1 0
Mode   Width  Height  BPP   BPSL   Write      FPS    Copy      FPS      Hz
 163     320     200    8    320   87.17  1428.22   79.63  1304.73   69.95
 164     320     240    8    320   85.10  1161.96   79.63  1087.27   67.28
 165     320     400    8    320   87.20   714.31   79.51   651.31   69.95
 166     320     480    8    320   85.07   580.76   79.49   542.68   67.28
 14F     400     300    8    400   78.06   682.10   75.51   659.84   73.41
 12D     512     384    8    512   78.43   418.28   75.48   402.56   81.61
 100     640     400    8    640   77.47   317.31   75.14   307.77   69.88
 101     640     480    8    640   77.11   263.20   74.98   255.92   60.16
 103     800     600    8    800   71.77   156.79   69.81   152.51   56.16
 105    1024     768    8   1024   65.32    87.10   47.24    62.99   87.08
 110     640     480   15   1280   62.28   106.29   57.59    98.29   60.05
 111     640     480   16   1280   62.28   106.29   57.59    98.29   60.05
 113     800     600   15   1600   51.41    56.16   41.85    45.71   56.16
 114     800     600   16   1600   51.41    56.16   41.85    45.71   56.15
 211     640     400   32   2560   39.04    39.97   26.54    27.17   70.10


Same as above, but with Pentium 75, MCLK /4 0 0
CPU:586 "GenuineIntel", Stepping ID:5, Model:2, TSC:Y, MSR:Y, MMX:n, CMOV:n
Mode   Width  Height  BPP   BPSL   Write      FPS    Copy      FPS      Hz
 163     320     200    8    320   70.53  1155.63   54.61   894.71   69.95
 164     320     240    8    320   66.82   912.32   54.61   745.61   67.28
 165     320     400    8    320   70.39   576.64   54.58   447.14   69.95
 166     320     480    8    320   66.66   455.08   54.39   371.32   67.28
 14F     400     300    8    400   64.63   564.78   54.61   477.22   73.42
 12D     512     384    8    512   65.06   347.00   54.28   289.47   81.61
 100     640     400    8    640   62.96   257.87   54.28   222.31   69.88
 101     640     480    8    640   62.54   213.46   54.28   185.27   60.16
 103     800     600    8    800   58.73   128.30   53.38   116.60   56.16
 105    1024     768    8   1024   65.32    87.10   40.96    54.61   87.11
 110     640     480   15   1280   54.04    92.23   45.13    77.02   60.05
 111     640     480   16   1280   54.05    92.24   45.13    77.02   60.05
 113     800     600   15   1600   51.41    56.16   35.62    38.91   56.16
 114     800     600   16   1600   51.41    56.16   35.60    38.89   56.16
 211     640     400   32   2560   35.10    35.94   22.06    22.59   70.10


Pentium 166 MHz (66.666*2.5) with Matrox Millenium 4MB
With option 'F', main mem sequential read speed 100.5 MB/s
Mode   Width  Height  BPP   BPSL   Write      FPS    Copy      FPS      Hz
 100     640     400    8    640   77.18   316.11   66.99*  274.39   70.17
 101     640     480    8    640   77.30   263.84   63.32*  216.12   60.02
 11D    1600    1200   16   3200   73.13    19.97   55.81    15.24   60.26
 11E    1600    1200   16   3200   73.13    19.97   55.81    15.24   60.26
 118    1024     768   32   4096   73.92    24.64   55.81    18.60   60.10
 119    1280    1024   16   2560   74.55    29.82   55.72    22.29   60.13
 11A    1280    1024   16   2560   74.55    29.82   55.71    22.28   60.14
 11C    1600    1200    8   1664   74.61    39.18   55.65    29.22   60.26
 107    1280    1024    8   1280   75.52    60.41   55.77    44.61   60.14
 116    1024     768   16   2048   75.52    50.35   55.75    37.17   60.10
 117    1024     768   16   2048   75.52    50.35   55.75    37.17   60.10
 115     800     600   32   3200   75.76    41.38   55.83    30.49   60.47
 114     800     600   16   1920   76.23    69.39   55.90    50.88   60.47
 113     800     600   16   1920   76.24    69.39   55.90    50.88   60.47
 112     640     480   32   2560   76.36    65.16   55.94    47.73   60.01
 105    1024     768    8   1024   76.83   102.44   55.96    74.62   60.10
 103     800     600    8   1024   76.89   131.23   55.96    95.51   60.47
 110     640     480   16   1280   76.99   131.40   55.96    95.51   60.01
 111     640     480   16   1280   76.99   131.39   55.96    95.51   60.02
*: other modes don't fit into the 256 kB L2 cache


AMD K6 208.333 MHz (83.333*2.5 MHz), 512 kB PB-cache, 64Mb EDO,
Asus P/I-P55T2P4 (HX chipset), Matrox Millenium 4Mb WRAM (41.666 MHz PCI-bus)
Main mem sequential read speed 125 MB/s. With option 'F'.
Mode   Width  Height  BPP   BPSL   Write      FPS    Copy      FPS      Hz
 100     640     400    8    640  116.78   478.35   84.60   346.51   70.18
 101     640     480    8    640  116.75   398.52   84.48   288.37   60.03
 103     800     600    8   1024  116.22   198.34   76.58   130.70   60.48
 105    1024     768    8   1024  115.84   154.45   68.73    91.64   60.11
 107    1280    1024    8   1280  113.58    90.86   62.94    50.35   60.15
 110     640     480   16   1280  115.99   197.95   76.56   130.66   60.03
 111     640     480   16   1280  115.99   197.95   76.56   130.66   60.03
 112     640     480   32   2560  115.27    98.37   62.93    53.70   60.03
 113     800     600   16   1920  114.73   104.43   62.94    57.29   60.48
 114     800     600   16   1920  114.74   104.44   62.93    57.28   60.48
 115     800     600   32   3200  113.54    62.01   62.93    34.37   60.48
 116    1024     768   16   2048  114.01    76.01   62.93    41.96   60.11
 117    1024     768   16   2048  114.02    76.01   62.93    41.96   60.11
 11C    1600    1200    8   1664  111.50    58.55   62.93    33.05   60.27
 118    1024     768   32   4096  110.99    37.00   62.93    20.98   60.11
 119    1280    1024   16   2560  111.59    44.64   62.92    25.17   60.15
 11A    1280    1024   16   2560  111.60    44.64   62.92    25.17   60.15
 11D    1600    1200   16   3200  108.44    29.61   62.93    17.18   60.27
 11E    1600    1200   16   3200  108.41    29.60   62.93    17.18   60.27


Same as above, but AMD K6 200 MHz (66.666*3 MHz, 33.333 MHz PCI-bus)
Mode   Width  Height  BPP   BPSL   Write      FPS    Copy      FPS      Hz
 100     640     400    8    640   93.78   384.13   71.92   294.57   70.18
 101     640     480    8    640   93.75   320.00   71.84   245.20   60.03
 103     800     600    8   1024   93.32   159.27   67.51   115.21   60.48
 105    1024     768    8   1024   93.06   124.08   62.99    83.99   60.11
 107    1280    1024    8   1280   91.96    73.56   59.45    47.56   60.15
 110     640     480   16   1280   93.44   159.47   67.53   115.26   60.03
 111     640     480   16   1280   93.44   159.47   67.53   115.26   60.03
 112     640     480   32   2560   92.88    79.26   59.45    50.73   60.03
 113     800     600   16   1920   92.73    84.41   59.45    54.12   60.48
 114     800     600   16   1920   92.73    84.41   59.45    54.12   60.48
 115     800     600   32   3200   91.30    49.86   59.46    32.47   60.48
 116    1024     768   16   2048   92.02    61.34   59.46    39.64   60.11
 117    1024     768   16   2048   92.01    61.34   59.45    39.64   60.11
 11C    1600    1200    8   1664   90.73    47.64   59.45    31.22   60.27
 118    1024     768   32   4096   90.90    30.30   59.46    19.82   60.11
 119    1280    1024   16   2560   90.79    36.32   59.46    23.78   60.15
 11A    1280    1024   16   2560   90.82    36.33   59.46    23.78   60.15
 11D    1600    1200   16   3200   89.27    24.38   59.46    16.24   60.27
 11E    1600    1200   16   3200   89.25    24.37   59.46    16.24   60.27


Same as above, but memory timings tweaked and 83.333*2.5 MHz (208.333 MHz)
Mode   Width  Height  BPP   BPSL   Write      FPS    Copy      FPS      Hz
 100     640     400    8    640  116.78   478.35   84.63   346.63   70.18
 101     640     480    8    640  116.76   398.54   84.55   288.59   60.03
 103     800     600    8   1024  116.22   198.35   79.20   135.17   60.48
 105    1024     768    8   1024  115.85   154.46   73.68    98.25   60.11
 107    1280    1024    8   1280  113.60    90.88   69.35    55.48   60.15
 110     640     480   16   1280  115.97   197.93   79.19   135.14   60.03
 111     640     480   16   1280  115.99   197.96   79.19   135.14   60.03
 112     640     480   32   2560  115.27    98.36   69.38    59.21   60.03
 113     800     600   16   1920  114.74   104.44   69.37    63.14   60.48
 114     800     600   16   1920  114.73   104.43   69.37    63.14   60.48
 115     800     600   32   3200  113.53    62.00   69.31    37.85   60.48
 116    1024     768   16   2048  114.01    76.00   69.33    46.22   60.11
 117    1024     768   16   2048  114.01    76.01   69.33    46.22   60.11
 11C    1600    1200    8   1664  111.52    58.56   69.24    36.36   60.27
 118    1024     768   32   4096  111.00    37.00   69.32    23.11   60.11
 119    1280    1024   16   2560  111.58    44.63   69.14    27.65   60.15
 11A    1280    1024   16   2560  111.59    44.63   69.14    27.66   60.15
 11D    1600    1200   16   3200  108.42    29.61   69.35    18.94   60.27
 11E    1600    1200   16   3200  108.44    29.61   69.35    18.94   60.27


Same as above, but with Cyrix 6x86-P200+
Mode   Width  Height  BPP   BPSL   Write      FPS    Copy      FPS      Hz
 100     640     400    8    640   98.46   403.31   92.98   380.85   70.18
 101     640     480    8    640   98.50   336.21   92.95   317.28   60.03
 103     800     600    8   1024   98.20   167.60   85.14   145.30   60.48
 105    1024     768    8   1024   97.98   130.64   77.18   102.91   60.11
 107    1280    1024    8   1280   96.85    77.48   71.05    56.84   60.15
 110     640     480   16   1280   98.16   167.53   85.14   145.31   60.03
 111     640     480   16   1280   98.16   167.52   85.15   145.32   60.03
 112     640     480   32   2560   97.56    83.25   71.17    60.73   60.03
 113     800     600   16   1920   97.49    88.74   71.14    64.75   60.48
 114     800     600   16   1920   97.49    88.74   71.14    64.75   60.48
 115     800     600   32   3200   96.59    52.75   70.98    38.76   60.48
 116    1024     768   16   2048   96.94    64.63   71.02    47.35   60.11
 117    1024     768   16   2048   96.94    64.63   71.02    47.35   60.11
 11C    1600    1200    8   1664   95.69    50.25   70.81    37.18   60.27
 118    1024     768   32   4096   95.65    31.88   70.97    23.66   60.11
 119    1280    1024   16   2560   95.62    38.25   70.65    28.26   60.15
 11A    1280    1024   16   2560   95.63    38.25   70.66    28.26   60.15
 11D    1600    1200   16   3200   94.44    25.79   71.00    19.39   60.27
 11E    1600    1200   16   3200   94.43    25.79   71.00    19.39   60.27


Same as above, but with Cyrix 6x86-P166+
Mode   Width  Height  BPP   BPSL   Write      FPS    Copy      FPS      Hz
 100     640     400    8    640   87.95   360.26   83.03   340.11   70.18
 101     640     480    8    640   87.97   300.27   83.05   283.46   60.03
 103     800     600    8   1024   88.05   150.27   76.13   129.94   60.48
 105    1024     768    8   1024   87.47   116.63   68.74    91.65   60.11
 107    1280    1024    8   1280   86.90    69.52   63.38    50.70   60.15
 110     640     480   16   1280   87.90   150.01   76.03   129.75   60.03
 111     640     480   16   1280   87.90   150.02   76.03   129.75   60.03
 112     640     480   32   2560   87.38    74.57   63.41    54.11   60.03
 113     800     600   16   1920   87.45    79.60   63.40    57.71   60.48
 114     800     600   16   1920   87.45    79.60   63.40    57.70   60.48
 115     800     600   32   3200   86.96    47.49   63.32    34.58   60.48
 116    1024     768   16   2048   86.79    57.86   63.37    42.25   60.11
 117    1024     768   16   2048   86.79    57.86   63.37    42.25   60.11
 11C    1600    1200    8   1664   86.19    45.26   63.26    33.22   60.27
 118    1024     768   32   4096   85.77    28.59   63.35    21.12   60.11
 119    1280    1024   16   2560   86.17    34.47   63.29    25.32   60.15
 11A    1280    1024   16   2560   86.17    34.47   63.29    25.32   60.15
 11D    1600    1200   16   3200   84.34    23.03   63.36    17.30   60.27
 11E    1600    1200   16   3200   84.33    23.03   63.36    17.30   60.27


Cyrix 6x86-P166+, no L2 cache, 32MB FPM ram, S3 Trio64V+ 2MB EDO (33.333 PCI)
Mode   Width  Height  BPP   BPSL   Write      FPS    Copy      FPS      Hz
 170     320     200    8    320   88.94  1457.21   50.39   825.67   69.94
 171     320     240    8    320   86.58  1182.11   50.36   687.58   67.27
 172     320     400    8    320   88.89   728.21   50.37   412.61   69.94
 173     320     480    8    320   86.52   590.66   50.36   343.81   67.27
 174     400     300    8    400   89.93   785.78   50.39   440.27   73.40
 175     512     384    8    512   91.06   485.66   50.35   268.56   81.60
 100     640     400    8    640   90.69   371.45   50.35   206.22   69.79
 101     640     480    8    640   88.15   300.89   50.35   171.85   72.84
 103     800     600    8    800   81.02   176.99   50.33   109.94   72.27
 105    1024     768    8   1024   70.90    94.53   50.34    67.12   69.92
 107    1280    1024    8   1280   59.73    47.79   50.22    40.17   60.20
 10D     320     200   15    640   91.21   747.17   50.37   412.63   69.78
 10E     320     200   16    640   91.21   747.21   50.37   412.63   69.78
 10F     320     200   32   1280   82.01   335.93   50.35   206.23   69.78
 110     640     480   15   1280   77.12   131.61   50.34    85.92   72.84
 111     640     480   16   1280   77.14   131.65   50.34    85.92   72.84
 112     640     480   32   2560   55.81    47.63   44.66    38.11   72.84
 113     800     600   15   1600   65.13    71.14   50.32    54.97   71.99
 114     800     600   16   1600   65.14    71.15   50.32    54.96   71.99
 115     800     600   32   3200   36.66    20.02   27.11    14.81   72.27
 116    1024     768   15   2048   45.08    30.06   36.83    24.55   70.04
 117    1024     768   16   2048   45.08    30.06   36.83    24.55   70.04
 120    1600    1200    8   1600   54.21    29.61   42.57    23.25   96.02



Windows tests run under Win95 with the newest S3 video driver @ 800x600x16bit.
(Cyrix 6x86-P120+, Trio64, 16 MB FPM, MG mobo with 256 kB cache)
Test 1: 59.96 MHz, 2-cycle EDO timings
Test 2: 85.01 MHz, FPM timings

                          Test 1     Test 2   Speedup
PC Labs Winbench v3.1   14644561   25862197     70.1%
Wintach v1.0: Text          53.4       79.8     49.4%
              Cad          236.4      275.9     16.7%
              Spreadsheet   57.5       89.2     55.1%
              Paint        105.2      143.3     36.2%


(Cyrix 6x86-P150+, Trio64, 16 MB FPM, Asus VX97 mobo with 512 kB cache)
Test 1: 59.96 MHz, 2-cycle EDO timings
Test 2: 85.01 MHz, FPM timings

                          Test 1     Test 2   Speedup
PC Labs Winbench v3.1   15587307   24906850     70.1%
Wintach v1.0: Text          54.1       81.7     51.0%
              Cad          262.4      311.8     18.8%
              Spreadsheet   58.0       89.6     54.5%
              Paint        106.7      144.8     35.7%



[8.] - OCTest

     OCTest (invoked by 'O' or 'Q'), is nice for verifying your computer's
   proper operation when you are overclocking or 'tweaking' your system.
   It uses 40 different block sizes. Code size for the main Random accessing
   routine (Part [5]) is 6.1 kB and 1.8 kB for the other parts, so total
   memory consumption [7.9 + test size + (1-4 depending on test size)] kB.
   Subtract 1.0 kB if using 'Q' instead of 'O'. Many CPUs have 8 kB code cache,
   so OCTest fits nicely in 8 kB.
     For Pentium, Pentium Pro and Cyrix 6x86, nice size to almost fully utilize
   L1 cache (and heat up the CPU most) might be 6 or 8 kB, for AMD K5, Pentium
   MMX and Pentium II 14 or 16 kB and for AMD K6 27 or 32 kB.
   Pentium Pro has L2 cache 'inside' the CPU, so using larger test sizes might
   heat it up more. Pentium II might also behave similarly. If you have
   more than one CPU, only the first one is tested. Safbench/Qsort32 do not
   work under WinNT. I'm not going to add SMP support (unless I get very
   rich...). If your computer works well in Win95 (one CPU in use) but fails
   every now and then in WinNT (two or more CPUs in use), your CPU might be
   overheating. PII CPUs can draw 40 watts or so, so take a good care of
   cooling. Also, the CPUs should be of the same stepping level. Or your
   power supply is crappy.

     When testing your new shiny mobo at 75/83/100 MHz, use 'O:M' option to
   verify memory/chipset operation. Press 'M' if you feel like pressing it,
   and F1 to F9 to make Safbench move mem like a crazy. I don't recommend
   testing CPU operation (when you are oveclocking) with Win95, since if
   you are unlucky, registry might get corrupted and you (might) have to
   reinstall Win95. Or FAT gets screwed up. Safbench just crashes if your
   CPU doesn't work OK and it doesn't modify the registry :) Run Safbench
   without multitasking. If your mobo/chipset isn't specked to work at
   75/83/100 MHz, don't assume that it works! At least do the following:
   use FULLNULL and NULLCHK under DOS to verify that HD/chipset works at
   that bus speed (try those progs for ~10 minutes). Test every HD!
   Then, run 'Safbench O' without multitasking and press 'H' (test both
   modes, fast and normal), test also with Use moves (press 'M' and F1-F9
   if you want), test the largest block size for ~10 minutes. If errors occur,
   slow down memory timings (in BIOS or with TweakBios), speculative
   leadoff(/read) causes problems in many configurations. If NULLCHK reports
   errors even with the lowest PIO modes, DON'T use the higher bus speed...

     OCTest was made because I couldn't find any program which actually
   verified that CPU+FPU work correctly. For example, playing Quake isn't
   necessarily a good way to verify that CPU+FPU work correctly, since you
   DON'T notice if there's ONE error in the calculations every
   1.000.000.000.000 CPU core cycles (means one error every hour on a
   266 MHz CPU). Of course assuming that the error is in the data dumped to
   the screen, not in the code, which might crash Quake or something...
    Quake doesn't verify CPU/FPU operation by doing the calculations twice.
   OCTest does. Also, some operating systems (Windows etc) are unstable even
   when the CPU is working correctly, so Win95 faults do not necessarily mean
   that your CPU/memory is malfunctioning. CPU just does what it's told to do
   and if it executes crappy code things will go wrong. Also, Winstone etc
   do not verify CPU/FPU/mem operation. One little error might happen once
   a hour without you noticing it. And it doesn't even use the whole memory
   space for testing.
    If your computer performs perfectly at 1.5x50 MHz with iP75 processor, it
   might not work at 1.5x66 MHz, BUT then again it might work at 2x60 MHz.
   Some crappy old (?) motherboards work well only at max 60 MHz rates (slow
   caches and such stuffs). And the same for new boards: all of them don't
   work at 75 or 83 MHz, see manufacturers' webpages or some 'tech sites', for
   example www.tomshardware.com.
    One more note: one 166 MHz CPU might overclock perfectly to 3.5x83.333=292
   MHz, but one other 166 MHz CPU only to 200 MHz. The only way to verify it
   is to test it!
    AMD K6/233 draws at maximum 28.3W at 3.2V, so every motherboard JUST
   doesn't work reliable with 233 MHz K6... At minimum 166-233 MHz K6 CPUs
   draw 1.0W (at stop clock mode) or 1.45-1.75W (at stop grant mode). For
   example, Linux puts the K6 in stop grant mode (with HLT-instruction) when
   it's 'idle', so the K6 runs/should run quite cool (I assume that K6 does
   the same as CPUs from 8086 and up do when HLT is executed: waits for an
   external interrupt, NMI or RESET). Then again, your BIOS reports the
   temperature while it's executing it's own instruction mix, I suppose it
   isn't designed to heat up the CPU (by intermixing FPU and so on). It might
   cause the K6 to dissipate only half of the maximum thermal power or
   something like that. SO, EVERYONE: could you tell me how to fetch the
   temperatures, for example from Asus mobos. Asus guys didn't manage to
   give me a reply about THAT (they replied, though).

     New key command while running OCTest, 'H', toggles HLT-instruction usage
   during part [5]. It puts your CPU in a kind of 'sleep mode'; it waits for
   external interrupt, NMI or RESET. You probably get only timer interrupts
   (18.2 Hz) and keyboard interrupts (hold down SHIFT, CTRL...) in OCTest.
   Press 'H' again and timer interrupts occur about ten times more often
   (you will see red 'F' (fast), gray 's' (slow) is the normal speed). Speed
   index will NOT be reliable when you've used this Halt-test.
   When HLT-instruction is being executed, you will see 'HALT:1'. Randomly
   also second/third HLT is executed, you will see 'HALT:2' and 'HALT:3'.
   There's a random delay between these 1-3 HLTs, it's 1-512 sub-loops of
   part [5]. I added this test so you can test whether or not your power
   supply transient response and decoupling capacitors are sufficient enough
   to handle the instantaneous current changes occurring during transitions
   from 'sleep mode' (or 'stop grant mode', or 'Auto HALT Powerdown state'
   or whatever) to full active mode. If you see errors occurring ONLY when
   HLTs are being executed, you might have shitty el-cheapo motherboard
   and/or voltage regulator. Errors in that case might be due to supply
   voltage going out of specified limits. If you have such a motherboard,
   I'd like to know about it (so I can make a list :) )!
   Linux executes HLT when it's idling, so it's also a good motherboard-test-
   proggy :)

     There's a smaller difference in speeds reported by the Speed Index at
   different block sizes than in Random accessing test in Safbench, because
   OCTest does a lot more accessing and FPU stuffs than Random accessing test.
     If you don't have any good main memory testing utils, try disabling L1
   and/or L2 cache(s) and run 'Safbench O:M'. It accesses memory more
   thoroughly than many of the (free and commercial) memory test programs
   I've seen. If you've got AMD K6, make sure it's rev. C or later; there's
   a bug which causes malfunctioning (in Linux, NT, some others) when it
   THINKS it detects selfmodifying code (this can only happen if you've got
   more than 32 MB memory).

     Cyrix 6x86-P150+ (120 MHz) got 100.000 at 8 kB block size (Asus VX97
   mobo with 512 kB cache, 16 MB FPM). There's no scaling for other test
   sizes.

   The test goes like this:
     1) initialize the whole ~8 MB memory array (or the size of free mem
        when testing block size number 40): writes random numbers, which
        are different every loop
     2) verify the array, if mismatch occurs, R/W error is reported
     3) [this part not used anymore]
     4) the block size to be tested is filled with random data
     5) memory is accessed (written, read, modified) randomly

        If Use moves=Y, the block size under testing is moved 64 bytes ahead
        using REP MOVSD instruction. CRC32 is calculated for the overwritten
        data and used later to assure good error detection. Data is checked
        before AND after the move, if there are differences, Move error is
        reported. Some (OLD) Pentium Pro motherboards have a bug which occurs
        when using "Fast string enable", it's a feature of the CPU
        which can be toggled in the BIOS/with a program. When it's on
        (default), REP MOVS moves 64 bits each clock cycle. If you have
        pressed function keys F1-F9, the move loop is repeated 10-90 times
        (you will see 'Use moves: [Y5]' if you press F5 and so on).

     6) CRC32 is calculated of the data made in step 5. If it differs from
        the previous loop's CRC32, Data error is reported. If CPU's registers
        weren't right, CPU error is reported.
     7) Data made in step 5 is QuickSorted
     8) Data is checked for validity; if they are not in the right order,
        Qso1 error is reported (for every 32b integer which is larger than
        the next integer)
     9) CRC32 of the QuickSorted data is calculated. If it differs from the
        previous CRC32, Qso2 error is reported.

        Then this whole process from 4 to 9 is executed for 15 seconds
        (or any time when 'O:xxx used). After that next block size will be
        tested. OCTest.$$$ file is made every 5 minutes (after completed loop),
        it just contains the screen capture.

     So: in the first loop from 4 to 9, CRC32 is calculated and CPU's
   and FPU's register values saved. If they don't match in the second
   loop, errors will occur on the screen. If a loop differs by more than
   two percents of the Speed index, 'Unstabiles' count is increased.
   This is not an error! The '#' character is a kind of speed-stability-meter;
   one position means ~.1% difference from the Speed index. The farther '#'
   is from the left side of the screen, the farther the elapsed loop time is
   from the average Speed index.

     Also FPU operation is checked: OCTest tests addition, multiply,
   dividing, square root and transcendentals. FDIV instruction is
   executed simultaneously with other CPU instructions so that your CPU
   heats up even more :) Next version might contain CPU's temperature
   display (if your mobo supports it). It takes about -3 minutes for the
   CPU to heat up to it's maximum temperature when running correct code
   stream designed specially to heat up that particular CPU (for example,
   MAX_POW for K6, but it isn't available (told me by AMD staff)).
   I scheduled FDIV operations so that they execute approximately all the
   time with CPU instructions.

     You will see flickering 'ESC', it's colour is changed every clock
   tick (18.2 times per second). It should flicker ALL THE TIME,
   if it doesn't, your CPU might have freezed. The flickering isn't just
   a stupid blink attribute. Note that OCTest's code might have 'hanged'
   even if the CPU executes the interrupt routine correctly. Caps Lock is
   blinked every second if 'L' option is used. If errors are detected, also
   the other two leds (Num+Scroll) start blinking. So you know what's
   happening even when the monitor is turned off.

     When overclocking, you will probably first get Qso2 errors (IF your
   CPU isn't working correctly), Data errors, CPU and Qso1 errors. FPU errors
   are rare, but can happen if FPU actually calculates erraneously or data
   cache's contents get screwed up. The last thing is probably crashing with
   Page Fault, General Protection Exception or Invalid Opcode (or Win-style
   crash with no error screens - just a stupid freeze). If you crash your
   system and see only black screen when you reset your computer, let it
   cool down for about 15 minutes and try again...
     When your CPU is overclocked, one thing which WILL most probably
   work incorrectly is CPU's CACHE: that's why QuickSort algorithm
   is used to test the cache: it stresses the caches (and of course
   main memory with big enough block sizes) and "shakes up" the caches pretty
   well (flushing/fetching the cache lines, marking them dirty/invalid...).
     Try disabling BTB, Write Gathering, WT_ALLOC or even CPU's (WB-)cache
   if your CPU isn't working very well. TweakBIOS v1.51a is a nice
   utility for changing memory timings etc. My 6x86-P120+ (100 MHz) works
   perfectly at 100 MHz, but at 120 MHz it's VERY unstable, if I don't
   cool the CPU with a hairdresser (with _COLD_ air) _OR_ use silicon jelly.
   It works without cooling if CPU's cache strategy is WT instead of WB.
   So you better get some silicon jelly (it's cheap). AMD K6/233 can
   dissipate up to 28.3 W thermal power, so good heatsink, fan and silicon
   jelly/grease are recommended. K6's case temperature must be below 70C or
   operation isn't guaranteed to be error-free. See AMD's WWW pages for a
   PDF-file which lists good sink+fan combinations. Remember also to take
   care of voltage regulator's cooling. Be careful when overclocking Cyrix
   6x86 pre v2.7, they generate a lot of heat and use a lot of power.
   I blew up my MG mobo (in 10 minutes, used OCTest) when I overclocked from
   100 to 120 MHz (my 6x86 is v2.5). Asus VX97 works OK.

     If your CPU STILL doesn't work OK (and you REALLY want to overclock it),
   try increasing core voltage. Higher signal level increases signal to noise
   ratio, thus decreases the possibility that transistors go into wrong states
   (causing crashes and/or malfunctioning). Also, CPU heats up more with higher
   voltage values so take a good care of cooling! For example, Asus TX97 mobos
   can monitor CPU's temperature (but unfortunately I haven't managed to get
   programming information from Asus about how to retrieve the temperatures).
   Don't blow up your CPU by giving it 5V or something...

     OCTest checks CPU's BTB, RSB, code/data TLB, pipelining (CPU's
   U/V-pipelines etc and CPU+FPU intermixing), memory, CPU's data cache,
   LOCK-signaling, segment loads and prefixes, FPU pipelining/transcendentals,
   everything. Everything but MMX. With Pentium P54C, amount of Data TLB misses
   start rapidly increasing from block size 213 kB and up (like someone cares).

     Press SPACE to skip to the next block size, also with 'O:M' (block
   size will be tested forever). Keypresses are checked after every
   completed loop, but at MAXIMUM 18.2 times/s. Exitlevel 0 is returned for
   flawless runs, 21 otherwise (if exited with ESC or autoexit).
   'Q' starts saving memory captures until an error occurs, OCTest is quitted
   then. 'W' is like 'Q', but saves all the allocated memory.
   'S' saves memory capture once, 'F' saves all the allocated memory once.
   'C' clears some stats, 'I' sets speed index to 100.000, 'D' does 'em both.
   'M' toggles memory move routine after Part [5] (it uses REP MOVSD to do
   it). F1-F9 selects move count, for example F4 executes the move loop 40
   times, F9 90 times.



[9.] - VESA info

   Mode      Hexadecimal VESA mode ('!' displayed after this if mode is not
             supported)
   Size      Width and Height of the video mode in pixels (graphics)
             or chars (text)
   Clrs      Number of different colours in the video mode
   BPP       Bits Per Pixel
   BPSL      Bytes Per Scan Line
   Size/B    Size of the video mode in bytes
   GCLB      G=graphics mode; Y=yes, -=no (text mode)
             C=color; Y=color, -=monochrome
             L=Linear Framebuffer supported for this mode, Y=yes, -=no
             B=Bank-switched mode supported for this mode, Y=yes, -=no
   Memtype   "    text" = text
             " CGA gfx" = CGA graphics
             " HGC gfx" = HGC graphics
             " 16c EGA" = 16-color (EGA) graphics
             "  packed" = packed pixel graphics
             "sequ 256" = sequ 256 (non-chain 4) graphics
             "  direct" = direct color (HiColor, 24-bit color)
             " YUV/YIQ" = YUV (luminance-chrominance, also called YIQ)
             "reserved" = reserved for VESA
             " OEM mem" = OEM memory models
   BO        Bios Output supported, Y=yes, -=no
D: Red       Red mask size/red field position
D: Grn       Green mask size/green field size
D: Blu       Blue mask size/green field size
D: Rsv       Reserved mask size/reserved mask position
D: P         Color ramp is programmable, Y=yes, -=no
D: A         Bytes in reserved field may be used by application, Y=yes, -=no
T: Char cell Char cell size in pixels

   D: only in Direct modes
   T: onlt in Text modes

   If reported VESA version is less than 2.0, '?' is displayed at the
   place of L and B, since they are officially available with VBE 2.0+.

   Yes, the output might seem 'messy', but at least it fits in 80 chars...



[10.] - Errorlevels

   0   No errors

   1   No FPU present

   2   Help screen displayed

   3   Not enough memory (needs 4112 kB)

   4   No 387+ FPU present (287 present)

   5   No VBE present

   6   Bad command line options

   7   No LFB modes found to test (try UniVbe or S3VBE20)

   8   Failed to map linear framebuffer

   9   Not enough memory to test video copy speeds

   10  Any Key Pressed

   12  Specified mode isn't graphical LFB mode (or it was invalid)

   13  VESA info displayed

   14  Could not set video mode

   15  CPL not 0 (needed for options 'B' and 'P'), try without multitasking

   16  CPUID and/or RDTSC not supported (for 'B' and 'P')

   17  Machine Specific Registers not supported ('B' and 'P' again)

   18  Under 8224 kB free memory, random access test skipped

   19  Under 8224 kB free memory, overclocking test skipped

   20  Specified more than 1000000 seconds for overclocking test

   21  Error(s) detected in OCTest, watch out not to burn your chips :)



[11.] - Final words

   If you encounter bugs (bugs? where?), don't hesitate to tell me.
   Also improvements accepted. See line #11 for more info.
   I reply to every message (sfarin@ratol.fi), so if you don't see a
   reply in a sensible time, try sfarin@hotmail.com!
   Thanks to Ralf Brown for Interrupt List 55. So long.
