Araara ḃ — Unicode
This is the fourth in a series of expository blog posts on Araara. The preceding post introduced the Aglobasa alphabet and the Akurto type that uses it. The series begins here.
## Alphabets are arbitrary
Before proceeding, I need to come clean. Choosing Aglobasa as our default alphabet is a convention. Its sounds make it convenient to read aloud, but the arithmetic does not depend on that particular choice.
**o ae bp cj dt fv gk hx iu lr mn sz wy**
But that’s kind of the point! I might prefer Aglobasa, you might prefer English, and someone else might want Greek letters. Once we specify the ordered alphabet and the rule assigning its values, we can interpret each other’s numerals.
The type name **Akurto** is simply a translation of **Ashort** into Globasa. By convention, I use Aglobasa for Akurto and ø-English for Ashort. Later, we will see how to describe alphabets themselves in Akurto. Our first step is to represent Unicode code points.
Before doing that, let’s consider another way of choosing one representation for each integer. We already have normal form and its zero-terminated version. Now we will define **short form**.
## Short form
Akurto normal form uses only positive letters for positive integers, and only negative letters for negative integers. Sometimes we could use fewer letters by allowing both signs.
For example, thirteen has normal form:
**ccba** = 5 + 5 + 2 + 1 = 13
But we could also write:
**de** = 14 - 1 = 13
We have gone a little beyond thirteen and then subtracted one. This takes two letters rather than four.
We choose the **short form** of an integer as follows:
1. Find the fewest letters needed to represent the integer, allowing both positive and negative letters. Among representations of that length:
1. If the normal form is among them, keep it.
2. Otherwise, arrange each candidate in reverse Aglobasa alphabetical order.
3. If several candidates remain, choose the alphabetically first, comparing from left to right in Aglobasa order.
The order here is always Aglobasa’s, not English’s. For example, both “gtpe” and “gtjb” represent 105 using four letters, and both are in reverse order. We choose “gtpe”, because the first difference is “p” versus “j”, and “p” comes first in Aglobasa. For zero, the short form is “o”.
## Unicode short form
Let us give our Unicode numerals a recognisable beginning: “u”. Since “u” has value −1094, the remaining letters must contribute **N + 1094** for the whole expression to sum to N.
For a Unicode code point N, define its **Unicode short form** (**usf**) by prepending “u” to the short form of N + 1094. We choose the suffix first, then add the prefix; the initial “u” stays in place.
For example, “ui” represents zero: −1094 + 1094 = 0. Adding eleven or thirteen gives:
**0** = **ui** (usf)
**11** = **uicca** (usf)
**13** = **uide** (usf)
## Akurto Unicode
A Unicode code point is an integer between zero and 1,114,111. It is usually written in hexadecimal with a ‘U+’ prefix. For example, uppercase ‘A’ has code point U+0041, or 65 in decimal. See the Unicode definition of a code point and the Basic Latin chart.
A visible character can consist of several code points; we encode those in order. For ‘A’, there is just one: 65. Its Unicode short-form suffix “ifdcc” sums to 1159; the initial “u” subtracts 1094, leaving 65:
'A' = **uifdcc**
## Unicode o-short form
To concatenate code points, we also need to mark where each numeral ends. Define **Unicode o-short form** (**uosf**) by appending “o” to its Unicode short form. Since “o” is zero, the numerical value stays the same:
**0** = **uio** (uosf)
**11** = **uiccao** (uosf)
**13** = **uideo** (uosf)
## Forming strings
We can now form strings by concatenating code points in Unicode o-short form.
For example, ‘Unicode’ has code points 85, 110, 105, 99, 111, 100, and 101. Encoding them one at a time gives:
‘Unicode’ = **uiffbao uigtbo uigtpeo uigttco uigjjeo uigvdco uiffdco**
The spaces are just for readability. Each encoded code point ends with “o”, and its short-form suffix contains no “o”: adding a zero would only make that suffix longer. We can therefore recover the boundaries even without spaces.
Even though Akurto has only twenty-five letters, we can express any sequence of Unicode code points. Isn’t that cool? We will call these **Akurto character code sequences** , or **Akccs**.
## Unicode normal form
Short form saves letters, but it is not our only convention. Define **Unicode normal form** (**unf**) by prepending “u” to the normal form of N + 1094. We can append “o” to this version too.
For ‘i’, whose code point is 105, Unicode normal form gives “uiffdcbb”; Unicode short form gives “uigtpe”. Both sum to 105. The latter simply uses fewer letters.
## What if we sum an Akccs?
As in the preceding posts, we can also evaluate the whole concatenated expression as one integer. For ‘Unicode’, the code points sum to 711:
**uiffbao uigtbo uigtpeo uigttco uigjjeo uigvdco uiffdco**
= **hhtj** = 365 + 365 − 14 − 5 = 711
We could consider “hhtj” a simple numerical signature, or a very basic hash, of ‘Unicode’. Given the string, its sum follows unambiguously. Changing how we represent the individual code points does not change that sum.
The reverse does not hold. Every anagram has the same sum, and even strings that are not anagrams can collide: ‘AC’ and ‘BB’ both sum to 132. The sum preserves neither order nor the individual characters. Still, it is a neat consequence of using the same arithmetic for numbers and text.
## Exercises
1. ‘B’ has code point 66. Find its Unicode short form, then encode the string ‘AB’ with terminating “o” letters.
2. Compare the encodings and sums of ‘AB’ and ‘BA’. What information survives concatenation, and what disappears when we take the sum?
We can now write integers, code points, and strings using Akurto. That brings us one step closer to describing the alphabets themselves.