Skip to content
jayconrod edited this page Jun 18, 2011 · 3 revisions

A type is a property of a value (or a definition that can be used as a value) that defines what operations can be applied to that value. You can also think of a type as a set of values. For instance, the Boolean type contains two values, true and false. A value of the Boolean type may be used in a conditional branch and can be passed to functions with Boolean parameters. However, you cannot perform arithmetic with a Boolean value since it is not a numeric.

Tungsten has many different categories of types, all of which are listed here, along with subtyping rules.

List of type categories

UnitType

The unit type has only one value (the unit value). unit is very similar to the void type used in languages based on C. Functions which have nothing useful to return will typically return unit. The reason we use unit instead of void is so we can say that all functions return values.

The unit value has zero size and provides no useful operations.

BooleanType

The boolean type contains the values true and false. A boolean value consumes 1 byte. boolean values can be used as part of conditional branches. They can also be used in some binary operations (AND, OR, XOR).

CharType

Deprecated.

StringType

Deprecated.

IntType

Integer types represent signed integers with a fixed bit width. There are four integer types: int8, int16, int32, and int64. Values of these types consume 1, 2, 4, and 8 bytes, respectively. Integer types support all binary and relational operators. There are also a number of instructions to cast integers to different sizes or to floating point types.

You can get a word-sized integer using the function IntType.wordType. This will return int64 or int32 depending on whether the module is 64-bit.

FloatType

Floating point types represent rational numbers of single or double precision. There are two float types: float32, and float64. Values of these types consume 4 and 8 bytes, respectively. Float types support all binary and relational operators. There are several instructions to cast floats to different sizes or to integer types.

PointerType

Pointer types represent addresses in memory. Pointer values are used to refer to other values indirectly. Every pointer type has an element type which is the type of value the pointer refers to. All pointers consume either 4 or 8 bytes, depending on whether the module is 64-bit. Currently, pointers only support EQUAL and NOT_EQUAL relational operators. No binary operators are supported. There are instructions for loading and storing values from pointers, as well as for performing pointer arithmetic.

Pointer types are written with a trailing *. For example, int64* is a pointer to a 64-bit integer.

NullType

The nulltype is a special pointer type which points to nothing. The null value is the only member of this type.

ArrayType

Array types represent fixed sized lists of elements, all of which have the same type. Array elements are stored contiguously in memory. The size of an array is the product of the number of elements and the size of each element.

Array types are written in brackets with a size and an element type. For example, [3 x int64] is the type of an array of 3 64-bit integers. This would consume 24 bytes of memory.

StructType

struct types are used to represent records: lists of values of different types. A [[struct definition | Struct]] determines the number of type of elements in struct value. The size of a record is determined by the struct definition. Elements are guaranteed to be stored in order, but they might not be in order; some alignment bytes may be added between elements for efficient loads and stores. The first element of a record will always be at offset 0 from the beginning of the record.

struct types are written with the keyword struct, followed by the name of the definition. For example, struct @Foo.

FunctionType

In Tungsten, functions can be treated as values. They can be called indirectly and passed as arguments. Like all other values, they must have types. A function type comprises a return type, a list of type parameter symbols, and a list of parameter types. Function values consume either 4 or 8 bytes depending on whether the module is 64-bit.

Function types are written in three parts. The list of type parameters comes first in brackets. The parameter type list follows, in parenthesis. The return type comes after an arrow. Both the type parameter and parameter type lists can be omitted if they are empty. For example, a function with no parameters or type parameters returning unit is written as any of the following:

->unit
()->unit
[]()->unit

A function taking two integer arguments returning an integer could be written as:

(int64, int64)->int64

The curried version of this function would be written as:

(int64)->(int64)->int64

A parameterized function which returns the same type as its argument would be written as:

[@f.T](type @f.T)->type @f.T

where @f.T is the name of the type parameter.

Object types

As in other languages, an object in Tungsten is an instance of a class. All objects are stored on the heap. Object types represent pointers to objects.

Due to subtyping, one object pointer may have many valid types. Each object has a class, and each class has a corresponding type. An object may also be typed by any of its superclasses or interfaces.

ClassType

Class types represent pointers to instances of a particular class. Class types can be used to make virtual method calls using the vcall instruction (the class's method list is used for this). They can also be used to load and store fields using the loadelement and storeelement instructions (the class's field list is used for this). It is possible to calculate the address of a field using the address instruction.

A class type comprises a class name and an optional list of type arguments. The number of type arguments must be the same as the number of type parameters for the classes, and each type parameter must be within the bounds of the corresponding parameter. Note that since upper and lower bounds of type parameters are always object types, type arguments must also be object types.

Here are some examples of class types:

class @Foo
class @Bar[class @A]
class @Baz[class @A, interface @B]

InterfaceType

Interface types represent pointers to objects which support a particular interface. Like class types, interface types can be used to make virtual method calls. However, the interface's method list is used, which makes it a slightly different operation. Interface types can also be used to load, store, and address fields with loadelement, storeelement, and address instructions, respectively. The field list of the interface's base class is used for this.

Here are some examples of interface types:

interface @Foo
interface @Bar[class @A]
interface @Baz[class @A, interface @B]

VariableType

A variable type is used to represent an unknown object type. Each variable type corresponds to a type parameter and is valid where that type parameter is in scope, i.e. within the class, interface, or function that declares the type parameter. A variable type is known to be a subtype of its type parameter's upper bound (if specified) and a supertype of the lower bound (if specified). It may be upcast to the upper bound. No other operations are valid for variable type values.

Variable types are written with the type keyword and the name of the type parameter. For example: type @f.X.

Subtyping rules

To understand subtyping, it helps to think of a type as a set of values. If we have two types, A and B, A is a subtype of B if and only if all the values in A are also in B. We denote the subtyping relation using the <: operator. So A <: B means A is a subtype of B.

The following rules are used to determine whether one type is a subtype of another:

  • S <: S (every type is a subtype of itself)
  • S <: T if S <: U and U <: T (the subtype relation is transitive)
  • forall T: nulltype <: T* (nulltype is a subtype of all pointer types)
  • [T11, ..., T1m](p11, ..., p1n)->r1 <: [T21, ..., T2m](p21, ..., p2n)->r2 if:
  • forall i in 1, ..., m: upper(T1i) = upper(T2i) and lower(T1i) = lower(T2i) (where upper and lower retrieve the upper and lower bounds of a type parameter)
  • forall i in 1, ..., n: expose(p2i) <: expose(p1i) (where expose replaces variable types with the upper bounds of the type parameters they represent)
  • expose(r1) <: expose(r2)
  • type @V <: T if upper(type @V) = T (a variable type is a subtype of its upper bound)
  • T <: type @V if lower(type @V) = T (a variable type is a supertype of its lower bound)
  • class @C[S1, ..., Sn] <: class @C[T1, ..., Tn] if:
  • forall i in 1, ..., n:
  • let Pi be the ith type parameter of class @C
  • if Pi is covariant: S1 <: T1
  • if Pi is contravariant: T1 <: S1
  • if Pi is invariant: S1 = T1
  • interface @I[S1, ..., Sn] <: interface @I[T1, ..., Tn] under the same rule as classes
  • class @C[T1, ..., Tn] <: D if D is inherited directly by class @C[T1, ..., Tn]
  • interface @I[T1, ..., Tn] <: D under same rule as classes

Note that Pointer and arrays types have no subtyping rule, even if the elements types they point to are subtypes. There is a good reason for this. Suppose B <: A and C <: A. If A* <: B*, then we could do the following (in pseudocode):

A* pa = ...
B* pb = pa
*pb = new C

Since pa and pb are aliases, they will both point to a value of type C. However, C is not a subtype of A, so pa would not be a valid pointer, and any operations on it would have undefined results.

Clone this wiki locally