Pebble Virtual Machine Systems Programmer's Handbook
By vantheman
Product of The Nuovi Orizzonti Company

The Pebble Virtual Machine is the first attempt at creating an implementation of the Expressive Instruction Set
Computing Architecture, also known as EISC. Pebble is a small interpreted virtual machine implementing what
should be considered a canonical example of what an EISC ISA should look like.

The EISC architecture is designed around the idea that operands should carry more than payloads, and that
addressing modes should be a modifier of the payload itself, potentially changing the entire instruction's
behavior, how the data is interpreted, and so on. However, what makes EISC different from CISC, is in the
EISC architecture, addressing modes are fully instruction-dependent, as well as each operand of an
instruction containing it's own addressing mode. 

In EISC, every addressing mode should have the ability to express an entire language construct. An example
constructed roughly from the Pebble ISA may be along the lines of

	STORE, 0, 0, 1 2, 8

In this, you have the instruction, being STORE. STORE takes a destination and a value, each with their own
addressing mode. However, the addressing modes themselves have the ability to express language concepts,
so in this case, addressing mode 0 for the first operand of STORE could mean to reference it as a pointer,
and the addressing mode of the payload may mean to resolve it as a mathematical addition statement. In this
case, say if store point 0 held some representation of store point 1, then the destination would not simply
be 0, but 0 would be used as a pointer referencing store point 1. Now if we look at the addressing mode for
operand 2, addressing mode 8 may mean to resolve as an addition statement. So this instruction may resolve
to storing the literal '3', being the result of the expression '1 2' using the addressing mode stating to
resolve it as an expression, to the location '1', which location '0' held some representation of and was
used as a pointer. However, since the addressing in itself should be uniformly applied, there is nothing
limiting the first operand of STORE from only being a pointer or a literal. For this example, addressing
mode 0 could mean to address as a pointer, then to resolve the expression held there. So if location 0
held say some representation of location 1, which then held the expression of say the addition of 1 and 2,
then this would resolve to storing the literal '3' to the store point of '3'. However, since addressing is
fully instruction-dependent, those addressing modes may mean something completely different for a different
instruction.

The main advantage of this is that it makes compilers targetting the language incredibly easy, as well as
a theoretical JIT compiler for the language itself much more optimal. Since high-level schematics are
kept in the bytecode language, the JIT does not have to worry about trying to re-construct them from lower
level primitives, allowing much more room for optimization inside of the virtual machine itself. However, 
this moves a lot of the complexity from the compiler into the virtual machine, which may not be such a bad
thing depending on how you look at it. However, the biggest advantage is that the machine is expressive and
uniform at the lowest level. Well, since Pebble is in itself a virtual machine, this is not the lowest level,
as it runs on some form of hardware with it's own ISA lower than this, but this is the lowest level that will
be discussed in this documentation. (well, besides for the compatability-related sections)

Pebble impliments an EISC ISA to be known as Pebble. Pebble is both the ISA and virtual machine. This means
that the Pebble textual bytecode language (the ISA) is executed inside of / by the Pebble virtual machine.
The Pebble textual bytecode language is not like most tradtional ISAs, instead implimenting higher level
language concepts into the bytecode itself. As stated above, this improves the potential room for optimzation.

The Pebble textual bytecode language maps up and down 1:1 to the internal Pebble bytecode language, which shall
not be discussed here, as that is treated simply as an implimentation detal. Pebble textual bytecode is not
stack-based or register-based, instead acting on named variables. There is a theoretical infinite number of
named variables, and the payloads of these variables do not have types exposed to the user. Said variables are
also of arbitrary size, whos size can be modified at any point in both directions, and can be modified at any
point.

The Pebble textual bytecode language is highly uniform and keyword-based. There are seven keywords with a few
addressing modes per keyword. As the EISC, Pebble's instructions have their own independent resolver with
each instruction having it's own instruction-dependent addressing modes.

These keywords include the New keyword, which assiagns a variable as the result of.. well something. It is 
generally described as more of an abstract 'do thing with data' concept. The syntax shall be the keyword
name, followed by a space, the destination, followed by another space, followed by the data payload and
a newline character. The destination payload supports syntax of <> to represent a pointer, which is a
variable which holds the name of another variable, "" to represent an expression, which has the variables
in the expression substituted to their values in-place and the result passed to the expression parser, {}
to represent a pointer, which the reference of said pointer should then be evaluated, '' to represent a
literal, or 'do not do anything with this data', and no markings to simply represent copying the data over.
An example of this could be

	New a '1' 	// a = 1
	New b a 	// b = copy of a (so 1)
	New a '1 + 2' 	// a now = literal string '1 + 2'
	New c {a} 	// c = result of a as an expression (so 3)
	New d <c> 	// d = pointer to c (so 'c') 
	
However, addressing modes can also be applied to the destination payload itself

	New e '8' 		// e = 8 - in the destination, no quotation means a literal, unline
		  		// in the data payload, where it means a copy of. Copying is not
		  		// supported for destination payloads, and well copy is specific to
		 		// the data payload of New, as it would not make sense anywhere else
		  		// really. I mean think about it - the destination is a copy of another
		  		// variable? does not make much sense.
	New f 'e' 		// f = 'e' - however this would have the same effect if it were
		  		// New f <e>, as New does not distinguish between a literal containing
		  		// a variable name and a pointer, as do most instructions.
	New <f> '2' 		// f as a pointer now = 2, and since 'f' contains 'e', e = 2.
	New calc 'a s++ b'	// store literal 'a s++ b'. The 's++' operator here means concat.
	New {calc} '3' 		// variable 'ab' (result of 'a s++ b') now = 3

The syntax for the destination payload of the New keyword is generally consitored to be the common case,
so for all future instructions specified here, the syntax for the destination payload of the New
instruction should be consitored the default and should be assumed to be the syntax unless specified
otherwise.

The Func keyword begins definition of a function. Every line of text until the End keyword will become
associated internally with the name of the function assiagned to the name of the function resolved
from the addressing mode. This function can later be called using the Call keyword, taking one
argument with the resolved name of the function to call after the addressing mode, and the If keyword,
which will call the function resolved from the addressing mode involving the first argument, if the
condition in the second argument resolved from the addressing mode results in 0. The second argument
differs from the standard, instead supporting the addressing schematics of the second argument of
the New instruction. Examples of the three may be

	New x 'y'	// sets variable 'x' to equal 'y'
	Func <x>	// create a function with the name 'y' (the result of the pointer to 'x')
	New z "2 + 1"	// create a variable named z with the result of '2 + 1', so '3'.
	End		// end the function definition
	Call <x>	// call the function with the name of 'y' (the result of the pointer to 'x')
	If <x> "1 < 2"	// call the function 'y' (result of hte pointer to 'x') if '1 < 2' is true
			// in this language, '0' is used as true, and all comparison statements return
			// '0' if true and '1' if false.

The Return keyword is used to return a function early. This will simply immediatly return to the caller
(old_pc + 1), just as If and Call, once it is seen and there is a function being ran.

The Escape keyword is used as the Pebble Foreign Function Interface (FFI). The Escape keyword should
simply take one argument, being the name of the host function to call resolved from the addressing
mode. Host functions are written in native Zig and compiled into the virtual machine itself. Once
the binary is compiled, host functions cannot be added or removed. Host functions only have access
to the virtual machine internal viarbale table (as in they can read and write to variables created
by the virtual machine and code running on it), but nothing more. Escape sequences (the host
functions) should follow the standard naming convention. Escape sequences should take arguments
using variables named by '__Escape_<name>_ARG<argnum>', where <name> is the name of the Escape
sequence being called and <argnum> is the argument number. Escapes should return into the
variables __Escape_<name>_RET<retnum>, where retnum is the return variable number, and should
only use internal variables __Escape_<name>_<varname>, where <varname> is the name of the
variable the Escape sequence wishes to define. This matches the convention used for
compiler-generated functions. This convention matches the Escape sequence naming convention, but
'Escape' should be replaced by 'Func' and other logical substitutions.

Escape sequences include

	std. 			- The standard library
	  - io. 		  - Basic input / output
	    - print 		    - Print character stream with a newline at the end
	    - printLn 		    - Print character stream without a newline at the end
	    - input		    - Blocking input grabbing
	  - term. 		  - Terminal manipulation libraries
	    - cursor. 		    - Cursor manipulation libraries
	      - moveAbsolute 	      - Move the cursor to an absolute position
	      - moveRelative 	      - Move the cursor to a relative position
	      - get 		      - Grab current cursor position
	    - color.		    - Terminal color manipulation libraries
	      - basic.		      - Terminal basic color manipulation libraries
	        - back		        - Set the basic background color of the terminal
		- fore			- Set the basic foreground color of the terminal
		- reset			- Reset the basic terminal color to the default

[TODO - finish this]

Bytecode Specifcations

The Pebble Bytecode format is a simple textualized format of the Pebble internal bytecode
format. This format should generally compile 1:1 to Pebble internal bytecode, as in
each instruction in this bytecode language should become one internal instruction, without
RISC-esc lowering. There are currently only two versions of the Pebble textual bytecode
format. These are simply versions one and two. Version one is rather depreciated, as it
was only ever used in the earliest version of the Pebble virtual machine, version Alpha 1.
Bytecode version two is the current standard of bytecode. Due to this, the schematics of
Bytecode version one shall not be described in this document

	Bytecode version two

		Pebble Bytecode version two has seven instructions. These include
		New, Func, End, If, Escape, Call, and Return. 

		The New instruction is used to store data into a variable. Variables
		have arbitrary English names. That is, any character included in the
		English alphabet is allowed, as well as the underscore (_), period (.), and
		minus (-) characters and numerics zero (0) through nine (9). The New
		instruction accepts two data payloads, no less. These include the
		first operand, being the destination payload, or the name of the
		variable to store to, and the seond operand, being the data payload,
		or the data to store to the variable. The data payload has a small
		bit of special syntax, known as an addressing mode, which allows
		a small bit of evaluation to be performed on it before resolution.
		The addressing modes are marked by the characters the payload is 
		wrapped in. These addressing mods include Literal (""), True Literal
		(''), Bare (), and Forced Evaluation ({}) (commonly shortened to
		Forced Eval or Indirect Eval). The Literal addressing mode accepts
		expressions, where each piece of text separated by a space is resolved
		as an item. An item can either be a literal, a variable, or a statement.
		All supported statements should be assumed to be statements, and all
		literals whos name matches the name of a currently established variable
		should be resolved as that variable. Essentially, all variables should
		be substituted to their value, all symbols referencing an operation
		should mark that operation, and all other values should be treated as
		a literal. The True Literal addressing mode should simply resolve to
		storing that data as plain data. The Forced Evaluation mode shoud
		resolve the variable given to it as a pointer, then resolve the 
		data in that variable as an expression. Operators follow standard
		ordering, but the ( ) operator syntax is explicitly excluded.
		Operators include addition '+', subtraction '-', multiplication '*',
		division '/', greater than numeric comparison '>', less than numeric
		comparison '<', equal to numeric comparison '==', not equal to numeric
		comparison '!=', string equal to comparison '?=', string ends with
		comparison 'e?=', string starts with comparison 's?=', and string
		contains comparison '-?='. All comparison operations should return
		0 if the operation is true, and 1 if the operation is false.
		On the second operand of this instruction, bare is used to represent
		a copy of a variable, while it is used to represent a variable on 
		the first operand.

		The Func instruction will begin recording a function to the function
		table. All of the instructions defined should be stored under the function
		name specified by the operand immediatly after the Func keyword. The
		operand can be any English text, and follows the same schematics as the
		New instruction's destination operand. This code can be executed later
		using the Call instruction, folowed by the name of the function to call
		using the same schematics, or the If keyword, which takes two operands,
		the first being the name of the function to call if the second operand
		as an expression following the syntax of the New instruction's Literal
		addressing mode resolves to 0 (True). The function should not be called
		if the expression does not resolve to 0. The expression should be wrapped
		in the '"' character, one being placed at the start, and one at the end.

		The Escape keyword is used to call a foreign function. Argument passing
		to the foreign function should be done via the variables named as
		__Escape_<escape name>_ARG<argument number>, and the Escape sequence
		should return to the variable __Escape_<escape name>_ARG<argument number>.
		The Escape is only allowed to modify variables starting with
		__Escape_<escape name>_ - this is used as a light enforcement boundary.
		The Escape instruction should take one operand as a literal, as the name
		of the Escape sequence to call. The name of the Escape sequence follows
		the same schematics as variable names.

		The Return instruction is used to return early from a function. This
		instruction simply returns to the caller and takes no operands.

		Comments can be used as either the '#' or the '/' keyword. Any line
		starting with a comment should be ignored. This also goes for lines
		which are invalid or not a known instruction.

		An example of a valid program may be

			# Store 5 into x
			New x '5'
			New y x
			/ Store a copy of x into y


			/ begin recording a function
			Func a
			Escape print
			/ Call the print Escape Sequence
			End
			/ End will automatically return from
			/ the instruction
			Call a
			/ Call function a
			If a "1 == 1"
			/ Call function a if 1 equals 1
			foo bar baz quax quux
			/ last line ignored since it is invalid
			/this comment is also valid since it is
			/ the first character
			#so is this one
[TODO - document AAE]

[TODO - finish this section]

[TODO - finish docs ToT]
