Skip to content

feat: Add predictive text and autocomplete support for Nuer (nus) - #353

Open
buayism wants to merge 3 commits into
keymanapp:masterfrom
buayism:buay_kun_gach_deng.nus.nuer
Open

buayism wants to merge 3 commits into
keymanapp:masterfrom
buayism:buay_kun_gach_deng.nus.nuer

Conversation

@buayism

@buayism buayism commented Sep 16, 2026

Copy link
Copy Markdown

- Closes keymanapp#345
- Integrates optimized book, natural sentence, and biblical wordlists
- Created and compiled by Buay Kun Gach Deng
@keyman-server

Copy link
Copy Markdown

Thank you for your pull request. You'll see a "build failed" message until the Keyman team has reviewed the pull request and manually initiated the build process.

Every change committed to this branch will become part of this pull request. When you have finished submitting files and are ready for the Keyman team to review this pull request, please post a "Ready for review" comment.

ɛn 276
raan 271
dee 263
/ci̱ 254

@darcywong00 darcywong00 Sep 17, 2026 •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Just wanted to confirm if the / character is supposed to be here. Generally we don't include punctuation characters in the wordlists.

There may 21 of these in the file.

And 3 more in natural_sentences_wordlist.tsv

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yes the / character is meant to be there. In Nuer language, when / character appears before a word like "/Cu", it means that word is negative. So "/Cu" and "Cu" are not the same. The one with a slash is negative and the word without a slash is positive

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

While technical, there is a way to resolve this: https://help.keyman.com/developer/18.0/guides/lexical-models/advanced/unicode-breaker-extension.

Categorizing the / as an "ALetter" type instead of its usual punctuation-style handling would let it be part of words. If care needs to be taken so that this only happens at the start of words and after spaces, that should still be possible, but would require greater fine-tuning.

@jahorton jahorton Sep 18, 2026 •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

To be clear, I haven't actually tested this within your model directly, but something like the following might be what you're looking for.

It definitely depends on if you use the / character in other ways, like how in English you might see and/or, which we interpret as and, /, or. The example below assumes this pattern exists in your language too, so please update us if that's a bad assumption!

  wordBreaker: (text) => {
    let customization = {

      /*** Definition of extra word-breaking rules ***/
      rules: [{
        match: (context) => {
          // When encountering [whitespace or start-of-text] + '/' , then a letter,
          // prevent wordbreaks between the '/' and the letter.  Allow them
          // in all other contexts.  (so, a/b would still split into a, /, b)
          if(context.propertyMatch(["WSegSpace", "sot"], ["Negation"], ["ALetter"], null)) {
            return true;
          } else {
            return false;
          }
        },
        breakIfMatch: false
      }],

      /*** Character class overrides for specific characters ***/
      propertyMapping: (char) => {
        if(char == '/') {
            return "Negation";
        } else {
          // The other characters already have useful word-breaking
          // property assignments.
          return null;
        }
      },

      /*** Declares any new, custom character classes to be recognized by the word-breaker ***/
      customProperties: ["Negation"]
    };

    /*** Connects all the pieces together for actual use ***/
    return wordBreakers['default'](text, customization);
  },

@buayism

buayism commented Sep 22, 2026

Copy link
Copy Markdown
Author

I edited the word breaker block so that / is treated as a negation prefix when it appears at the start of the word or immediately after whitespace. i've built the model successfully and i've confirmed it keeps "/cu" as one token

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

feat: Add predictive text and autocomplete support for Nuer

4 participants