=====================================================
VE COMPRESSED 32-BIT EXECUTABLE FORMAT (VETools 1.21)
=====================================================

------------------------------------
Creating and using VE executables
------------------------------------

Watcom's Linker cannot directly link to VE format.  Link to OS/2 LE format
as usual, and then use the LE2VE tool to create a VE file:

	le2ve vp.exe vp.ve

This VE file can then be run directly with VERUN or it can be appended to
VELOADER.EXE to create a PMODE/W executable or VELOD4GW.EXE to create a
DOS/4GW executable.  If you want an executable:

	le2ve -s vp.exe vp_ve.exe

will give you a PMODE/W version, and -S will give you a DOS/4GW version.
Note that VPLOADER.EXE and VPLOD4GW.EXE must be in the same directory as
le2ve for this to work.  If you don't want "loading" messages to come up
during decompression:

	le2ve -Q vp.exe vp_ve.exe

If you already have a VE file without a stub, it can be easily made into
an executable with COPY:

	copy /b veloader.exe+vp.ve vp_ve.exe

Note that DOS/4GW Pro is not currently supported as VELoader can't currently
wind its way past DOS/4GWPro to get to the LE header.  You can use LE2VE to
compress a 4GWPro program, though.

------------------------------------
Comparison to other formats
------------------------------------

The only Watcom executable compression I know of is PMWLITE, the compressor
included with PMODE/W V1.33.  Here are compression results for selected
executables (mostly games):

		PMWBind	   PMWLite	   VE (-i1 -r1)	   VE (-i2 -r2)
1. STTNG.OVL:	1,781,679  639,453(64.1%)  602,931(66.2%)  518,854(70.9%)
2. DUKE3D.EXE:	  962,239  531,418(44.8%)  465,784(48.4%)  420,212(56.3%)
3. DOOM.EXE:	  564,505  313,259(44.5%)  297,307(47.3%)  264,316(53.2%)
4. ASSAULT.EXE:	  214,184  104,344(51.3%)   92,997(56.6%)   83,860(60.8%)
5. RA2.EXE:	  647,649  343,400(47.0%)  334,755(48.3%)  303,474(53.1%)
6. VP.EXE:	  473,392  272,359(42.5%)  255,407(46.0%)  229,550(51.5%)
7. WAR2.EXE:	  661,619  335,280(49.3%)  317,426(52.0%)  284,592(57.0%)
8. MEGAD.EXE:	1,723,385  518,574(69.9%)  229,785(86.7%)  165,784(90.4%)

A) The executables were as follows:
   1. Star Trek: The Next Generation, A Final Unity.  Main executable file.
   2. Duke Nukem 3D V1.3d.
   3. Ultimate DOOM.
   4. Rebel Assault V1.0. Main executable file.
   5. Rebel Assault II. Main executable file.
   6. VGAPaint 386 V1.3beta (version 4.705).
   7. Warcraft II V1.22 retail.
   8. MegaDrive 0.0.5 emulator.  Selected to cream the opposition. :-)

B) "PMWBind" represents the executable after the 16-bit stub was replaced
   by PMODE/W V1.33.  This shrunk the size of some of the executables
   significantly since they were prepended with DOS/4GW Professional, which
   takes up ~250K.  It actually expanded others slightly since PMW1.33 is
   slightly larger than the Watcom DOS/4GW stub.

C) All executables ended up with PMODE/W as the stub after compression (-s
   used with LE2VE).

D) VELoader works fine with DOS/4GW; PMWLITE executables can only be run
   under PMODE/W.  Also, VELoader adds 4k to the VE executable size, while
   PMWLITE support is built into PMODE/W.

------------------------------------
Background
------------------------------------

The OS/2 LE format used by DOS/4GW and PMODE/W is nice but it leaves a lot
to be desired.  The format was originally designed to allow (1) dynamic
load-time linking of external shared libraries directly to calls in the
program, and (2) demand loading directly from the executable without having
to load, relocate, and swap the image out to virtual memory first.  However,
it is extremely inefficient for DOS/4GW programs, for the following reasons:

1) DOS/4GW does not support DLL linking.  Fields in the header for library
   linking are unnecessary.  This is a miniscule penalty, however.
2) DOS/4GW does not demand load even if virtual memory is enabled.
3) To allow demand loading, the relocation data holds the full source offset
   to modify as well as the target offset to apply the relocation toward,
   completely ignoring the data already present in the image at those
   addresses.  Most relocation data is 32-bit, meaning that most relocations
   require from seven to nine bytes:
	* 2 bytes for type and flags
	* 2 bytes for intrapage offset
	* 1 byte for object number
	* 2 or 4 bytes for target offset

4) The format allows for multiple segments even though all segments are
   loaded sequentially into a single memory block and the image is stored
   in that fashion.  This requires extra relocation information for no
   reason.

The VE format is designed to solve these problems.  It has the following
advantages:

1) SIMPLE header format.  And in a sensible order.

2) Much more efficient relocation compression.  Unlike the LE relocation
   format, the VE format uses a sorted and delta'd form of relocs in addition
   to using the relocated areas within the executable image to reduce the
   space required for each reloc from 7-9 bytes to approximately 1.1 bytes.
   This format change alone makes VE reloc data smaller than LE reloc
   data by anywhere from 8:1 in most cases to 10:1 in some cases.

   The most dramatic case, however, is with the MegaDrive 0.0.5 emulator.
   It has 90,578 32-bit relocations which occupy 810,627 bytes in the
   executable, almost half the executable.  After initial relocation
   sorting, the relocation block size drops to 90,803 bytes, about 8.9:1.
   After LZH compression, the size further drops to 6,699 bytes, which
   is a ratio of 121:1 over the LE data!

3) Image and relocation data can be compressed beforehand and decompressed
   upon loading.  Image data typically compresses from 1.8:1 to 2.5:1 while
   relocation data often compresses between 2:1 and 6:1.

------------------------------------
Why it works
------------------------------------

The LZSS and LZH compression methods used in the VE format is well known.
LZH (Lempel-Ziv-Huffman) is the basis for Deflate, the current compression
method for the popular Info-Zip, PKZIP, and GNU gzip tools, as well as
the format for the new PNG image format.  Although the LZH method used in
the VE format is actually a slightly simplified version designed for
efficient decompression, this is not where the bulk of the saved space
comes from.

Where VE makes a killing is in the relocation data.  Relocation data in LE
executables has two faults that kill compression ratios:

1) Unnecessary data.  This data is compressed by LZH, but it still takes
   a significant amount of space.  For instance, few, if any, DOS/4GW apps
   have any relocations other than 32-bit offset.  Eliminating the type
   tags alone eliminates one byte per relocation, and deleting the usually
   zero flags byte eliminates another byte.

2) Ascending addresses.  LE relocations have a page offset.  The changing
   addresses cannot be compressed by LZ, both because the offset is only
   two bytes and because the values never repeat, and the addresses kill
   the Huffman compression, because the low byte varies around the entire
   8-bit space.

VE uses simple solutions to cut the relocation data size:

1) Lift the redundant data out.  For instance, all 32-bit relocations can
   be encompassed by a wrapper; all relocations within that wrapper are
   known to be 32-bit and thus all the 32-bit relocation tags can be
   discarded.  This change cuts about a fourth of the relocation data out.

2) Treat the executable as a single block.  There is no reason to segment
   the executable, because fragmentation is not a problem, the memory space
   is not shared (or paging will be on) and the blocks will be page
   aligned anyway.  Now all the object codes can be cut out, slicing
   an eighth off the relocation block.

3) Use the space in the executable.  The LE format requires target addresses
   to be placed in the relocation data.  VE places this data in the image
   itself, where space exists unused in the LE format.  Bye bye to one
   half of the data.

4) Delta the relocations and use MIDI-like encoding to reduce small deltas
   to one byte.  This now eliminates one more eighth out of the block,
   reducing the average size of a relocation to little more than one byte.

The last solution is the most critical, because it also immensely increases
the efficiency of the LZSS/LZH compression.  Large pointer tables tend to
balloon executable sizes, because with 13 bytes per entry (9 per relocation
plus 4 per pointer) almost none of it is compressible.  Emulators are
particularly prone to this.  The relocation data for the table under VE,
however, would look like this:

00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 ...

LZH has a field day on this kind of data!  Under the VE implementation,
maximum compression is 332:1.


------------------------------------
Header
------------------------------------

+00h	"VE"	Identification tag
+02h	0000h	Version
+04h	long	Offset to image
+08h	long	Offset to relocs
+0ch	long	Offset to library list (unused)
+10h	long	Offset to export list (unused)
+14h	long	Length of image
+18h	long	Length of relocs
+1ch	long	Length of library list (unused)
+20h	long	Length of export list (unused)
+24h	long	Uncompressed image length
+28h	long	Uncompressed relocs length
+2ch	long	BSS length
+30h	long	Flags
+34h	long	Entry point
+38h	long	Stack point
+3ch	long	Offset to resources (reserved)
+40h	byte	Image compression
+41h	byte	Relocs compression
+42h	byte	Library list compression (unused)
+43h	byte	Export list compression (unused)

* All offsets are from the start of the VE header.
* All variables are in LSB to MSB order.
* Compression values:
	0	None
	1	LZSS
	2	LZH

--------------------------------------
Relocs
--------------------------------------
00h	End of relocs
01h	32-bit reloc
02h	48-bit (far) reloc
03h	16-bit (segment) reloc
04h/05h	32-bit reloc to external library by name/number (unused)
06h/07h	48-bit reloc to external library by name/number (unused)
08h/09h	16-bit reloc to external library by name/number (unused)

[for 48-bit and 16-bit relocs, a segment of 0 is code and 1 is data]


+00h	byte	Reloc type
	[word]	Library number for reloc types 04h-09h
	[word]	Function name reference ID for reloc types 04h, 06h, 08h
	[word]	Library function for reloc types 05h, 07h, 09h
+01h	word	Number of relocs (zero means 64k relocs)
+03h	long	First reloc
+07h	byte	Reloc differentials (number of relocs - 1)
	00-7fh	Next reloc is <size> to <size>+127 bytes away
	80-ffh	(a) add 128-16384 bytes to the next value [last value]
		(b) << 7
		(c) << 7
		(d) << 7

--------------------------------------
Export list (unused)
--------------------------------------
00h	long	Address of first export
04h	long	Number of exports

	+00h	long	Delta 
	+04h	byte	Length of name; may be zero
	+05h	byte	Name

--------------------------------------
Library list (unused)
--------------------------------------
00h	long	Number of referenced libraries

	+00h	byte	Length of name; may be zero
	+01h	byte	Name

??	long	Number of name references

	+00h	byte	Length of name; may be zero
	+01h	byte	Name

--------------------------------------
Secondary compression: LZSS
--------------------------------------

+00h	long	Number of codes
+01h		Code groups:

	+00h	byte	Bit flags; bit 0 is first code in group
			0: code is a literal byte
			1: code is a reference to earlier data

	For each code in the group:
		If the code is a literal, copy the next byte
		If the code is a reference, the next short indicates the ref:
			bits 0-3 is length-3
			bits 4-15 is start of reference: current-4096+offset

Note: If the reference ends past the current point data should be repeated.
      This is automatic if a forward move is used in a buffer.

--------------------------------------
Secondary compression: LZ-Huffman
--------------------------------------

Distance encoding
-----------------
Offset	   Extra bits
   1	         0
   2		 0
   3		 0
   4		 0
   5-6		 1
   7-8		 1
   9-12		 2
  13-16		 2
  17-24		 3
  25-32		 3
  33-48		 4
  49-64		 4
  65-96		 5
  97-128	 5
 129-192	 6
 193-256	 6
 257-384	 7
 385-512	 7
 513-768	 8
 769-1024	 8
1025-1536	 9
1537-2048	 9
2049-3072	10
3073-4096	10

Length encoding
---------------
Len	Xbits
 3	  0
 4	  0
 5	  0
 6	  0
 7	  0
 8	  0
 9	  0
10	  0
11-12	  1
13-14	  1
15-16	  1
17-18	  1
19-22	  2
23-26	  2
27-30	  2
31-34	  2
35-42	  3
43-50	  3
51-58	  3
59-66	  3
67-82	  4
83	  0

Literal/Distance tree: 256+24 = 280 codes
Length tree:			 22 codes

Format:
1. Bit lengths for L/D bit length tree, packed into nibbles, bits 0-3 before
   bits 4-7 (16 nibbles).
2. Bit lengths for length tree, packed into nibbles the same way (22 nibbles).
3. Literal/distance bit-length tree (280 codes).
4. Encoded data.

Bit packing order:
All values are packed MSB to LSB; this is the reverse of the way the Deflate
method packs Huffman codes.

--------------------------------------
Applying relocs
--------------------------------------
1) Load and depack relocs if necessary.
2) Process blocks sequentially:
  a) 32-bit relocs: add base address to <target> (long).
     48-bit relocs: add base address to <target> (long).
		    place CS in <target+4> if <target+4> is 0 or DS if
		    <target+4> is 1 (target+4 is short).
     16-bit relocs: place CS in <target> if <target> is 0 or DS if <target>
		    is 1.
  b) Set <target> to the base address plus <first reloc>.
  c) All relocs in the same block are ascending in memory by delta offsets.
     The smallest delta offset is 0, which indicates the next immediate
     relocation item for that type of reloc (2 bytes for 16-bit, 4 bytes
     for 32-bit, and 6 bytes for 48-bit).  Thus, for 32-bit relocs, the
     range of deltas possible with one delta byte is 4-131.
  d) If the delta is too large to fit into one byte, the delta is split into
     multiple bytes with 7-bits of the delta in each byte.  The bytes are
     ordered from most significant to least, with all bytes but the last
     having the MSB set.  Thus the string 81828374 is equivalent to the
     delta (0x01<<28) + (0x02<<14) + (0x03<<7) + (0x74), or 0x2081F4.
  e) There is no delta for the first reloc.  A block with two relocs would
     only have one delta.

-------------------------------------
Generating relocs
-------------------------------------
1) Multiple blocks with the same reloc type are permitted but it is
   recommended that this only occur if all relocs of that type cannot fit
   into a single block.  The multiple blocks should not share relocation
   areas (i.e. relocations in one block should not lie within 10000h-20000h
   while another block has relocations from 18000h-28000h), in the interest
   of greater compression.
2) All relocs of a particular type should be sorted in ascending address
   order to minimize the number of extension delta bytes (>80h) required
   as well as maximize compression derived from any secondary compression
   such as LZSS.

-------------------------------------
Loading VE executables
-------------------------------------
0) Save ES; it contains the PSP segment.  The program name can be obtained
   from here if necessary.
1) Load in VE header.  If the VE executable is attached to a preloader
   such as VELoader it is necessary to read the MZ header and then find
   the end of the LE executable to find the VE header.
2) Allocate memory according to the VE header for the image and relocs.
   BSS memory should immediately follow the image and thus both should
   be allocated as a single block.
3) Read in image and decompress.
4) Read in relocs and decompress.
5) Apply relocs to the image.  CS and DS applied to image for 16-bit and
   48-bit relocs should be the same as those supplied to the program at
   execution start.
6) Free relocs memory.
7) Zero BSS memory.
8) Switch stack to offset noted in VE header.  This is necessary since
   the Watcom library requires the stack pointer to be at the end of
   the BSS segment for stack checking.
9) Load ES with the PSP segment as obtained at the start of the loader.
10) Near call the program.
11) INT 21h AX=4C00h if the program happens to return.

-------------------------
ENGLISH MAJOR EXAMINATION
-------------------------

The VE 32-bit Executable Format and the tools LE2VE, VERUN, VELOADER, and
VELOD4GW are Copyright (C) 1995-1997 by Avery Lee.  All Rights Reserved.
The VE 32-bit Executable Format and the toole LE2VE, VERUN, VELOADER, and
VELOD4GW are free software; you may distribute and/or modify them under
the terms of the GNU General Public License.  It should have been included
in the archive(s); if not, write to the Free Software Foundation, Inc.,
675 Mass Ave., Cambridge, MA 02139, USA.

Any trademarks, registered trademarks, etc. not specifically mentioned as
being so are recognized here.  No offense or infringement is intentional.

Any patent and/or copyright violations in this package are certainly
unintentional.

