unicode_blocr

Version, currently 0.1.02 versions

github.com/j8r/unicode_blocr

Identify the Unicode block to which a character belongs

2 stars
0 dependents
License: ISC

Nothing has been indexed for 0.1.0 yet. The tag is recorded, its shard.yml has not been read, so the manifest and dependency list below are empty because they are unknown rather than because they are absent.

Installation

# Add this to your shard.yml
dependencies:
  unicode_blocr:
    github: j8r/unicode_blocr
    version: ~> 0.1.0

Then run:

shards install

shard.yml

No shard.yml has been indexed for 0.1.0. You can read it on the repository.

Dependencies

Unknown: the shard.yml for this version has not been read yet.

README

This README is the one indexed from the repository at its latest ref, not from the tag for this version.

Unicode Blocr

Build Status ISC

Identify the Unicode block to which a character belongs.

This library can be used to identify the type of characters, and eventually filter them (for example emoticons).

In Unicode, a block is defined as one contiguous range of code points (https://en.wikipedia.org/wiki/Unicode_block).

Installation

Add the dependency to your shard.yml:

dependencies:
  unicode_blocr:
    github: j8r/unicode_blocr

Usage examples

Basic

To print the block range to which the character belongs:

require "unicode_blocr"

puts UnicodeBlock.new 'a' #=> UnicodeBlock::BasicLatin
puts UnicodeBlock.new 'é' #=> UnicodeBlock::Latin1Supplement

Filter characters

To keep all characters inferior to a block range, here MiscellaneousSymbolsandPictographs and Emoticons, we delete all characters belonging to blocks above MiscellaneousSymbolsandPictographs.

require "unicode_blocr"

puts "hi😊".delete &.ord.>= UnicodeBlock::EnclosedIdeographicSupplement.value #=> hi
puts "café".delete &.ord.>= UnicodeBlock::BasicLatin.value #=> caf

License

Copyright (c) 2018-2019 Julien Reichardt - ISC License