ISO/IEC TR 15285:1998
(Main)Information technology — An operational model for characters and glyphs
Information technology — An operational model for characters and glyphs
The purpose of this Technical Report is to provide a general framework for discussing characters and glyphs. The framework is applicable to a variety of coded character sets and glyph-identification schemes. For illustration, this Technical Report uses examples from characters coded in ISO/IEC 10646 and glyphs registered according to ISO/IEC 10036. This Technical Report -differentiates between coded characters and registered glyphs -identifies the domain of use of coded characters and glyph identifiers -provides a conceptual framework for the formatting and presentation of coded character data using glyph identifiers and glyph representations This Technical Report describes idealized principles that were not completely followed in coding characters for ISO/IEC 10646 and in registering glyphs according to ISO/IEC 10036. The fact that ISO/IEC 10646, ISO/IEC 10036, and other standards do not completely follow the principles in the model does not invalidate the model and does not diminish the utility of having the model.
Technologies de l'information — Modèle pour l'utilisation de caractères graphiques et de glyphes
General Information
- Status
- Withdrawn
- Publication Date
- 19-Dec-1998
- Withdrawal Date
- 19-Dec-1998
- Technical Committee
- ISO/IEC JTC 1/SC 2 - Coded character sets
- Drafting Committee
- ISO/IEC JTC 1/SC 2 - Coded character sets
- Current Stage
- 9599 - Withdrawal of International Standard
- Start Date
- 08-Sep-2021
- Completion Date
- 12-Feb-2026
Get Certified
Connect with accredited certification bodies for this standard

BSI Group
BSI (British Standards Institution) is the business standards company that helps organizations make excellence a habit.

NYCE
Mexican standards and certification body.
Sponsored listings
Frequently Asked Questions
ISO/IEC TR 15285:1998 is a technical report published by the International Organization for Standardization (ISO). Its full title is "Information technology — An operational model for characters and glyphs". This standard covers: The purpose of this Technical Report is to provide a general framework for discussing characters and glyphs. The framework is applicable to a variety of coded character sets and glyph-identification schemes. For illustration, this Technical Report uses examples from characters coded in ISO/IEC 10646 and glyphs registered according to ISO/IEC 10036. This Technical Report -differentiates between coded characters and registered glyphs -identifies the domain of use of coded characters and glyph identifiers -provides a conceptual framework for the formatting and presentation of coded character data using glyph identifiers and glyph representations This Technical Report describes idealized principles that were not completely followed in coding characters for ISO/IEC 10646 and in registering glyphs according to ISO/IEC 10036. The fact that ISO/IEC 10646, ISO/IEC 10036, and other standards do not completely follow the principles in the model does not invalidate the model and does not diminish the utility of having the model.
The purpose of this Technical Report is to provide a general framework for discussing characters and glyphs. The framework is applicable to a variety of coded character sets and glyph-identification schemes. For illustration, this Technical Report uses examples from characters coded in ISO/IEC 10646 and glyphs registered according to ISO/IEC 10036. This Technical Report -differentiates between coded characters and registered glyphs -identifies the domain of use of coded characters and glyph identifiers -provides a conceptual framework for the formatting and presentation of coded character data using glyph identifiers and glyph representations This Technical Report describes idealized principles that were not completely followed in coding characters for ISO/IEC 10646 and in registering glyphs according to ISO/IEC 10036. The fact that ISO/IEC 10646, ISO/IEC 10036, and other standards do not completely follow the principles in the model does not invalidate the model and does not diminish the utility of having the model.
ISO/IEC TR 15285:1998 is classified under the following ICS (International Classification for Standards) categories: 35.240.30 - IT applications in information, documentation and publishing. The ICS classification helps identify the subject area and facilitates finding related standards.
ISO/IEC TR 15285:1998 is available in PDF format for immediate download after purchase. The document can be added to your cart and obtained through the secure checkout process. Digital delivery ensures instant access to the complete standard document.
Standards Content (Sample)
�
TECHNICAL ISO/IEC
REPORT TR 15285
First edition
1998-12-15
Information technology — An operational
model for characters and glyphs
Technologies de l’information — Modèle pour l'utilisation de caractères
graphiques et de glyphes
Reference number
��
�
ISO/IEC TR 15285: 1998 (E) © ISO/IEC
Contents
Page
Foreword. iii
Introduction . iv
1 Scope . 1
2 References. 1
3 Definitions . 1
4 Character and glyph distinctions. 2
5 Operational model . 3
6 Glyph selection . 6
7 Summary. 8
Annex A: Bibliography . 9
Annex B: Characters . 10
Annex C: Glyphs . 14
Annex D: Font models. 17
Annex E: Examples of character-to-glyph mapping . 22
Annex F: Recommendations of the original report. 24
© ISO/IEC 1998
All rights reserved. Unless otherwise specified, no part of this publication may be reproduced or
utilized in any form or by any means, electronic or mechanical, including photocopying and
microfilm, without permission in writing from the publisher.
ISO/IEC Copyright Office ` Case Postale 56 ` CH-1211 Genève 20 ` Switzerland
Printed in Switzerland
�
© ISO/IEC ISO/IEC TR 15285: 1998 (E)
Foreword
ISO (the International Organization for Standardization) and IEC (the
International Electrotechnical Commission) form the specialized sys-
tem for worldwide standardization. National bodies that are members
of ISO or IEC participate in the development of International Stan-
dards through technical committees established by the respective
organization to deal with particular fields of technical activity. ISO and
IEC technical committees collaborate in fields of mutual interest.
Other international organizations, governmental and non-govern-
mental, in liaison with ISO and IEC, also take part in the work.
The main task of a technical committee is to prepare International
Standards, but in exceptional circumstances a technical committee
may propose the publication of a Technical Report of one of the fol-
lowing types:
†� Type 1, when the required support cannot be obtained for the
publication of an International Standard, despite repeated efforts;
†� Type 2, when the subject is still under technical development or
where for any other reason there is the future but not immediate
possibility of an agreement on an International Standard;
†� Type 3, when a technical committee has collected data of a dif-
ferent kind from that which is normally published as an Interna-
tional Standard (“state of the art”, for example).
Technical Reports of types 1 and 2 are subject to review within three
years of publication to decide whether they can be transformed into
International Standards. Technical Reports of type 3 do not necessar-
ily have to be reviewed until the data they provide are considered to
be no longer valid or useful.
ISO/IEC TR 15285, which is a Technical Report of type 3, was pre-
pared by Joint Technical Committee ISO/IEC JTC 1, Information
technology, Subcommittee SC 2, Coded character sets, and Sub-
committee SC 18, Document processing and related communication
(which has since been reorganized into SC 34, Document description
and processing languages).
iii
ISO/IEC TR 15285: 1998 (E) © ISO/IEC
Introduction
People interpret the meaning of a written sentence by the shapes of
the characters contained in it. For the characters themselves, people
consider the information content of a character inseparable from its
printed image. Information technology, in contrast, makes a distinc-
tion between the concepts of a character’s meaning (the information
content) and its shape (the presentation image). Information technol-
ogy uses the term character (or coded character) for the information
content, and the term glyph for the presentation image. A conflict ex-
ists because people consider characters and glyphs equivalent.
Moreover, this conflict has led to misunderstanding and confusion.
This Technical Report provides a framework for relating characters
and glyphs to resolve the conflict because successful processing and
printing of character information on computers requires an under-
standing of the appropriate use of characters and glyphs.
Historically, ISO/IEC JTC 1/SC 2 has had responsibility for the devel-
opment of coded character set standards such as ISO/IEC 10646 for
the digital representation of letters, ideographs, digits, symbols, etc.
ISO/IEC JTC 1/SC 18 has had responsibility for the development of
standards for document processing, which presents the characters
coded by SC 2. SC 18 standards include the font standard, ISO/IEC
9541, and the glyph registration standard, ISO/IEC 10036. The Asso-
ciation for Font Information Interchange (AFII) maintains the 10036
glyph registry on behalf of ISO.
This Technical Report is written for a reader who is familiar with the
work of SC 2 and SC 18. Readers without this background should
first read Annex B, “Characters”, and Annex C, “Glyphs”.
This edition of the Technical Report does not fully develop the com-
plex issues associated with the Chinese, Japanese, Korean, and
Vietnamese ideographic characters used in East Asia. In addition,
although it discusses the process of rendering digital character infor-
mation for display and printing, it avoids discussing the inverse proc-
ess of character recognition (that is, converting printed text into char-
acter information in the computer).
iv
TECHNICAL REPORT © ISO/IEC ISO/IEC TR 15285:1998 (E)
Information technology —
An operational model for characters and glyphs
ISO/IEC 10646-1: 1993, Information tech-
1 Scope
nology — Universal Multiple-Octet Coded
Character Set (UCS) — Part 1: Architecture
The purpose of this Technical Report is to
and Basic Multilingual Plane.
provide a general framework for discussing
characters and glyphs. The framework is
applicable to a variety of coded character 3 Definitions
sets and glyph-identification schemes. For
For the purpose of this Technical Report,
illustration, this Technical Report uses ex-
the following definitions apply. The defini-
amples from characters coded in ISO/IEC
tions have been extracted from the ISO/IEC
10646 and glyphs registered according to
9541-1: 1991 and ISO/IEC 10646-1: 1993
ISO/IEC 10036.
standards.
This Technical Report
3.1 character: A member of a set of ele-
†� differentiates between coded charac- ments used for the organisation, control, or
ters and registered glyphs representation of data. (ISO/IEC 10646-1:
1993)
†� identifies the domain of use of coded
characters and glyph identifiers 3.2 coded character set: A set of unam-
biguous rules that establishes a character
†� provides a conceptual framework for
set and the relationship between the char-
the formatting and presentation of
acters of the set and their coded represen-
coded character data using glyph iden-
tation. (ISO/IEC 10646-1: 1993)
tifiers and glyph representations
3.3 font: A collection of glyph images
This Technical Report describes idealized
having the same basic design, e.g. Courier
principles that were not completely followed
Bold Oblique. (ISO/IEC 9541-1: 1991)
in coding characters for ISO/IEC 10646 and
in registering glyphs according to ISO/IEC 3.4 font resource: A collection of glyph
10036. The fact that ISO/IEC 10646, representations together with descriptive
ISO/IEC 10036, and other standards do not and font metric information which are rele-
completely follow the principles in the model vant to the collection of glyph representa-
does not invalidate the model and does not tions as a whole. (ISO/IEC 9541-1: 1991)
diminish the utility of having the model.
3.5 glyph: A recognizable abstract
graphic symbol which is independent of any
2 References
specific design. (ISO/IEC 9541-1: 1991)
ISO/IEC 9541-1: 1991, Information technol-
3.6 glyph collection: An identified set of
ogy — Font information interchange — Part
glyphs. (ISO/IEC 9541-1: 1991)
1: Architecture.
3.7 glyph image: An image of a glyph, as
ISO/IEC 10036: 1996, Information technol-
obtained from a glyph representation dis-
ogy — Font information interchange — Pro-
played on a presentation surface. (ISO/IEC
cedures for registration of font-related iden-
9541-1: 1991) [See the definition of graphic
tifiers.
symbol.]
ISO/IEC 10180: 1995, Information technol-
3.8 glyph metrics: The set of information
ogy — Processing languages — Standard
in a glyph representation used for defining
Page Description Language (SPDL).
ISO/IEC TR 15285: 1998 (E) © ISO/IEC
the dimensions and positioning of the glyph independently and contain terminology that
shape. (ISO/IEC 9541-1: 1991) requires explanation.
3.9 glyph representation: The glyph In information technology, characters are
shape and glyph metrics associated with a abstract information elements in the domain
specific glyph in a font resource. (ISO/IEC of coding for data representation and, in
9541-1: 1991) particular, data interchange. Coded charac-
ter set standards assign numeric values,
3.10 glyph shape: The set of information
character names, and representative (sam-
in a glyph representation used for defining
ple) images to each character contained in
the shape which represents the glyph.
a coded character set. Typically a character
(ISO/IEC 9541-1: 1991)
is given a name, which also serves to dif-
ferentiate it from the other characters of the
3.11 graphic character: A character, other
coded character set. The precise semantics
than a control function, that has a visual
and appearance of the information elements
representation normally handwritten,
in any given implementation are not defined
printed, or displayed. (ISO/IEC 10646-1:
by those standards for coded character
1993)
sets. This apparent lack of definition is not
considered to be a defect in the standards.
3.12 graphic symbol: The visual repre-
Recognizing that the information may be
sentation of a graphic character or of a
acted upon (deciphered, sorted, trans-
composite sequence. (ISO/IEC 10646-1:
formed, formatted, archived, presented,
1993) [See the definition of glyph image.]
etc.) by many different application proc-
esses during its lifetime, standards for
3.13 presentation [of a graphic symbol]:
coded character sets are defined as a basis
The process of writing, printing, or display-
for information interchange.
ing a graphic symbol. (ISO/IEC 10646-1:
1993)
In information technology, glyphs are ab-
stract presentation elements in the domain
3.14 presentation form: In the presenta-
of presentation processing. The ISO/IEC
tion of some scripts, a form of a graphic
10036 standard for glyph registration de-
symbol representing a character that de-
fines the process for assigning glyph identi-
pends on the position of the character rela-
fiers, glyph descriptions, and representative
tive to other characters. (ISO/IEC 10646-1:
(sample) images to each glyph submitted
1993)
for registration. The precise usage and ap-
pearance of these presentation elements in
3.15 presentation surface: A virtual rep-
any implemented font resource is not de-
resentation of a presentation medium
fined by those glyph registration activities.
(page, graphic display, etc.) maintained by
As with the coded character set standards,
the presentation process, on which all glyph
this apparent lack of definition is not con-
shapes are to be imaged. (ISO/IEC 9541-1:
sidered to be a defect in the standards.
1991)
Glyph identifiers are unambiguously as-
signed as a basis for tagging presentation
3.16 repertoire: A specified set of charac-
elements in and among interchanged font
ters that are represented in a coded charac-
resources, recognizing that the font-specific
ter set. (ISO/IEC 10646-1: 1993)
design information may vary from one font
resource to another.
4 Character and glyph
distinctions Characters and glyphs are closely related,
with many attributes in common and yet
The character and glyph definitions in
with distinctions that make it essential that
clause 3, which were taken from ISO/IEC
they be managed in information processing
10646 and ISO/IEC 9541, were developed
as separate entities. The ISO/IEC 10646
standard recognizes the distinction between
© ISO/IEC ISO/IEC TR 15285: 1998 (E)
characters and their visual representation †� A glyph conveys distinctions in form or
by defining the term, graphic symbol. The appearance. A glyph has no intrinsic
graphic symbol of SC 2 standards and the meaning.
glyph image of SC 18 standards represent
†� One or more characters may be de-
equivalent concepts. However, glyph and its
picted by no, one, or multiple glyph rep-
associated ISO/IEC 9541 terminology are
resentations (instances of an abstract
preferred when referring to presentation and
glyph) in a way that may depend on the
presentation processing.
context.
The historical association of characters and
glyphs has resulted in character sets main-
5 Operational model
taining distinctions that cannot be founded
on distinctions in meaning, but only on dis-
5.1 Character and glyph domains
tinctions in shape. Similarly, the glyph regis-
tration authority and the SC 18 font re-
Character information has two primary do-
source model have made use of criteria
mains as illustrated in Figure 1 on the next
based on meaning to abstract potential dis-
page. The first pertains to the processing of
tinctions in shape. In practice, ISO/IEC
the content, that is, the meaning or phonetic
10646 contains characters that appear to be
value of the character information. This is
instances of glyphs, while the glyph registry
depicted on the left side of the figure. The
prescribed by ISO/IEC 10036 contains
second pertains to the presentation of the
glyphs that appear to be designated as ab-
content of the character information. This is
2)
stract characters. In both cases, the ideal
depicted on the right side of the figure.
nature of characters and glyphs has been
Each domain places different requirements
compromised to a degree. For example, in
on the representation of the character in-
ISO/IEC 10646-1, SC 2 coded the “¿” glyph
formation. For example, searching for char-
into the character U+FB01 LATIN SMALL
acter information in a database and sorting
LIGATURE FI “ ” for round-trip integrity with
records containing character information
1)
other standards. (See Annex B.5 The
entail different requirements from those
“round-trip rule”.) Also, the JTC 1 Registra-
found in presenting characters on paper.
tion Authority (AFII) for ISO/IEC 10036
The former processes are primarily con-
could have registered the same glyph iden-
cerned with the content of data and have
tifier for the “$” glyph and used it for the
little or no concern about the appearance
U+0041 LATIN CAPITAL LETTER A “$” charac-
that the data may take.
ter, for the U+0391 GREEK CAPITAL LETTER
ALPHA “ ” character, and the U+0410 On the other hand, a composition and lay-
CYRILLIC CAPITAL LETTER A “ ” character. out process has little concern for the con-
However, AFII instead registered three tent of data, but great concern about its
glyph identifiers. appearance. In general, processing of char-
acter information in the content domain is
Within the realm of information technology,
independent of font resources, whereas
an ideal characterization of characters and
processing in the presentation domain is
glyphs and their relationship may be stated
strongly dependent on the font resource
as follows:
used for the presentation of the character
information. However, processes that per-
†� A character conveys distinctions in
form transformations from one domain to
meaning or sounds. A character has no
the other are aware of both the content and
intrinsic appearance.
appearance of characters. For example, a
character recognition process converts im-
�����������������������������������������������������������
�����������������������������������������������������������
1) This Technical Report describes a character in
terms of its 10646 code position (U+FB01), its 2) ISO/IEC 6429 also depicts a 2-layer structure. For
10646 name (LATIN SMALL LIGATURE FI), and illus- ISO/IEC 6429, the data layer could use charac-
trates it with a representative glyph in quotation ters, and the presentation layer could use glyphs
marks (“ ”). to present the characters in the data layer.
ISO/IEC TR 15285: 1998 (E) © ISO/IEC
\D/RXW
*\OQRKS
&KDUDFWHUV *O\SKV
WWXQLLD
QWQWH&RFHQD$SHDUS
VQHU2SRF3UJQVVLRFH3U
QHZWHEHRLUDS2
QLVPD’R
Q(WWDP)RU
DU\D
R26UDUW
RN&
QJNLFHKUD&PDP*UWQFRLOHH6R0
�
Figure 1 — Character and glyph domains
ages into coded characters. Also, a pDUD� The recognition that two separate domains
JUDSK�OHYHO�K\SKHQDWLRQ�SURFHVV�LV�DQ�H[� of processing are commonly applied to
DPSOH�RI�D�OD\RXW�SURFHVV�WKDW�UHTXLUHV�FRQ� character-based information leads to a con-
WHQW�LQIRUPDWLRQ� clusion that two primary forms of this infor-
mation are needed:
It is not possible, in general, to code data in
such a way as to optimize one process 1. a content-oriented form that is amena-
without reducing the performance of other ble to immediate content-based proc-
processes. Even within the content domain, esses and that can be easily converted
the nature of the character coding employed to and from other optimized forms
for textual data affects the type or types of
2. an appearance-oriented form that facili-
processing to be performed on the data; no
tates imaging of content
single coding can optimize more than a few
such potential processes. Given this situa-
These are, respectively, the character-
tion, the best solution is to formulate an
based form and the glyph-based form. Fail-
independent, logical character coding that,
ure to recognize this distinction between the
when necessary, can be transformed into
character domain and the glyph domain has
another coding more amenable to the proc-
led to the development of inconsistent stan-
essing required. For example, in the case of
dards and inconsistent systems that lack
searching, character data is often recast
functional separation of the two domains.
into specific forms that facilitate quick
searches. For sorting, a specially created
5.2 Composition, layout, and
sort key is required. In addition, because
presentation
ISO/IEC 10646 contains glyph-like charac-
ters, it is expected that implementations
As depicted in Figure 2 on the next page,
may choose to canonicalize or normalize
the composition and layout process (for
such characters by translating them to nor-
glyph selection and positioning) spans both
mative characters. A presentation subsys-
processing domains. If attention is restricted
tem that employs such a technique may
to the text portion of this process, the pres-
require that character data be normalized
entation of character-based information
prior to presentation.
requires three primary operations:
XVH
6SHOOKHF LQJ
5HFRQ WLJQL
3ULQ &K DFWH
UWUGHU
’LVSO
6H FK
’D WD U\
HU WLRQV 2SHDW QV
HVVLQJ
DWLR
G6XEVWRQ
6HOHFWL
© ISO/IEC ISO/IEC TR 15285: 1998 (E)
&RQWHQW�EDVHG�
3URFHVVLQJ
OOH
FK
VRS
HUW
%RWK3RUQL
,Q
\JOSKHQLHUVLGWIL
$SSHDUDQFH�\DOVSL
EDVHG�
3URFHVVLQJ
�
Figure 2 — Composition, layout, and presentation
3)
†� selecting the glyph representations many. This is particularly true for ISO/IEC
needed to display character data 10646 implementation level 3, which uses
combining characters. In its fully general
†� positioning the glyph shapes on the
form, the relationship is a context-sensitive
presentation surface
M-to-N mapping where M > 0, N � 0. For
some characters in ISO/IEC 10646-1, for
†� imaging the glyph shapes
example, the U+FEFF ZERO WIDTH NO-
BREAK SPACE character, no glyph (N=0) is
Glyph selection is the process of selecting
defined.
(possibly through several iterations) the
most appropriate glyph identifier or combi-
The SC 18 document-processing model
nation of glyph identifiers to render a coded
separates the glyph selection and layout
character or composite sequence of coded
operations from the operation of imaging
characters. Coded characters and their as-
the glyph shape to permit document inter-
sociated implicit or explicit formatting infor-
change between the processes. Glyph se-
mation (for example, specification of the
lection and positioning are part of the com-
font and its size) represent the primary in-
position and layout process, whereas
puts to composition and layout processing,
imaging the glyph shape is part of the pres-
and glyph identifiers (or the associated
entation process. The result of composition
glyph metrics and glyph shapes) represent
and layout is a final-form document, which
the primary output from composition and
contains font identifiers, glyph identifiers,
layout processing. The degree of glyph se-
and coordinate positions, along with either
lection sophistication varies widely among
references to font resources or the actual
existing standards and implementations.
font resources themselves. Such a docu-
ment form contains all the necessary infor-
The relationship between coded characters
mation required to present the formatted
and glyph identifiers may be one-to-one,
�����������������������������������������������������������
one-to-many, many-to-one, or many-to-
3) The necessity for mapping characters to glyphs
(glyph selection), not its complexity, is one of the
motivations for developing this operational model
for characters and glyphs.
3ULQ WLQJ
’LQJ
IRUPDWLRQ
QWDWHVH
RXW/D\
(Q WU\
IRUPDWLRQ ,Q
DUDF&K ’DWD
&RP LW LRQ
HFNLQJ
6S
6HDUFK LQJ
6RUWL QJ
ISO/IEC TR 15285: 1998 (E) © ISO/IEC
document on some presentation medium. being formatted). For example, German
An example of such a final form document text could use the “˜” and “‡” glyphs
is an SPDL (ISO/IEC 10180) document for quotation marks; and French text,
instance. the “'” and ““” glyphs.
An important aspect of this document- †� When the U+002D HYPHEN-MINUS “�”
processing model is that it begins with character is encountered, a composi-
coded-character data as its input and pro- tion and layout process may have to
duces either glyph-based data or directly determine if it is used in a mathematical
imaged glyph shapes as its output. That is, formula, as a separator between figures
it incorporates a transformation from a (digits), as a separator between words,
coded-character representation of a docu- or as a separator between syllables.
ment’s content to a glyph-based coding of a Depending on which context applies, it
document’s appearance. The latter may will select a minus sign, a figure dash, a
only be visible to the internal mechanisms quotation dash, or a hyphen dash (or
of an operating system or a user-interface possibly a hyphen point) glyph to dis-
subsystem in the case that the result is di- play the character.
rectly imaged for presentation. However,
NOTE: Because the ISO/IEC 10646 repertoire
even these systems frequently support
includes the necessary characters, some appli-
some form of output that contains the glyph-
cations resolve quotation marks and the hyphen-
based final form of the document.
minus illustrated in the previous two points by
converting to the appropriate 10646 characters
as they are input rather than selecting the ap-
6 Glyph selection
propriate glyphs for presentation.
While some earlier formatting systems as-
†� When a parenthesis or square bracket
sume a one-to-one correspondence be-
character is encountered in a document
tween characters and glyphs, this is inade-
being formatted in vertical lines (for ex-
quate for many applications and scripts.
ample, with East Asian ideographs), a
Many contemporary composition and layout
composition and layout process may
systems support more complex glyph-
need to choose a vertical variant glyph
selection processes that provide for the
form of the parenthesis or square
representation of sequences of multiple
bracket. It may also perform a similar
character codes by a single glyph or by the
selection for certain other characters
use of sequences of glyphs to represent
such as U+30FC KATAKANA-HIRAGANA
certain characters. In general, glyph selec-
PROLONGED SOUND MARK “ ”, U+2014
tion needs to be based on style information
EM DASH “‡”, U+2025 TWO DOT LEADER
and context as well as on the character data
“���”, etc.
itself. For example, consider the following:
†� When an Arabic letter is encountered in
†� When the U+0022 QUOTATION MARK “�”
an Arabic, Farsi, Urdu, etc. document, if
character is encountered, a composi-
the Arabic style being used to display
tion and layout process may need to
the text is of the Simplified Naskh type,
determine whether it begins or ends a
a composition and layout process may
quotation and then choose either an
have to choose an isolated, initial, me-
opening or closing quotation mark
dial, or final glyph form for the given let-
glyph (“‡” or “·”) as appropriate. In ad-
ter according to its context in the
dition, the process may select glyphs
document. For example, glyphs for
depending on the language of the text
U+0647 ARABIC LETTER HEH “ ” are
being formatted (or the formatting style
shown in Figure 3 on the next page.
specifications that apply to the content
© ISO/IEC ISO/IEC TR 15285: 1998 (E)
the choices required to determine an ap-
propriate glyph are based solely on (1) the
context of a character within a document,
� � � �
(2) the style specifications that apply to a
given character, or (3) a combination of the
Isolated Initial Medial Final
context and style specification. All of the
choices required for the examples shown
Figure 3 — Glyphs for ARABIC LETTER HEH
above fall into one of these categories.
However, in general, glyph selection can
†� In addition, Arabic typography makes
only be made as an integral part of the en-
extensive use of ligatures. For exam-
tire composition and layout process. Con-
ple, Figure 4 shows the isolated forms
sider the following:
of U+0627 ARABIC LETTER ALEF “ ” and
U+0644 ARABIC LETTER LAM “ ”, and
†� When hyphenating a line of text during
then the two ligature forms used when
composition, a composition and layout
Lam is followed by Alef.
process may insert a hyphen glyph
form at the end of a line if the line is
broken at a hyphenation point.
� � � � †� If hyphenating a German text between
the letters “F” and “N”, a composition
Alef Lam Ligature Ligature and layout process may replace the “F”
Lam-Alef Lam-Alef
with a “N”.
Isolated Final
†� If during the composition of a German
Figure 4 — Two example ligatures in an
text, the character sequence “III” is en-
Arabic font
countered, a composition and layout
process may select two distinct (non-
†� When a U+0930 DEVANAGARI LETTER RA
ligated) glyph forms for U+0066 LATIN
“‰” is encountered in a Hindi, Marathi,
SMALL LETTER F “I”. However, if the po-
Sanskrit, etc. document, a composition
sition for a hyphen (a hyphen point)
and layout process may have to deter-
should occur before the last “I”, that is,
mine whether a subscript, superscript,
at “II I”, then a composition and layout
half (“eyelash”), or full form glyph is re-
process may select an ff ligature glyph
quired according to context. If a sub-
“j”, followed by a hyphen (on the first
script form is required, a composition
line), and begin the subsequent line
and layout process may have to choose
with a normal glyph for the third and
from one of a number of possible sub-
final “I”.
script forms depending on the glyph to
which it is to be attached. Figure 5
†� A composition and layout process may
shows an example of this.
select small cap glyph forms for the first
line of a paragraph of Roman text.
†� A composition and layout process may
‰ … ˙ � h
select a swash glyph form for the first
Full Super- Subscripts Half
and last character of each line of a
script
paragraph.
Figure 5 — Glyphs for DEVANAGARI LETTER
†� A composition and layout process may
RA
select one of a number of possible
variant glyph forms for certain Arabic
The process of glyph selection is some-
letters depending on whether more or
times implemented as a separate part of
less space is available for composing a
composition and layout because many of
line of Arabic text.
ISO/IEC TR 15285: 1998 (E) © ISO/IEC
†� When justifying a line of Arabic text, a mation technology distinguishes two re-
composition and layout process may lated, but distinct, domains:
start by selecting ligature glyph forms
–� The processing domain uses
that consume the smallest amount of
coded characters to represent the
linear space in a line, and then sequen-
character’s meaning.
tially replace these ligatures with com-
–� The presentation domain uses
ponent ligatures or component non-
glyph identifiers to represent the
ligature glyphs such that more linear
character’s image.
line space is consumed up to the re-
quired line measure. Alternatively, a
†� Processes are available to convert be-
composition and layout process may
tween the two domains:
start justification by selecting no liga-
tures and then sequentially select liga-
–� Presentation processing takes the
tures that consume a smaller amount of
coded-character data plus any
linear space until the desired line
formatting data plus font informa-
measure is achieved or until an inter-
tion to display and print character
word space stretch threshold is
data.
reached (that is, a point at which inter-
–� A character recognition process
word spaces can be stretched to justify
scans images, analyzes the
the line to the desired measure).
shapes, and outputs the coded
characters that correspond to the
In summary, the glyph-selection process is
shapes.
primarily applicable to behavior occurring at
the end or beginning of individual lines of
†� Depending on the script and the par-
text, or within the context of justifying or
ticular font or fonts used, glyph selec-
altering the measure of a given line during
tion can be straightforward or relatively
line composition. A system supporting the
complex.
capabilities illustrated in the preceding ex-
amples must include glyph selection as an
–� It is straightforward when a one-to-
integral part of the composition and layout
one correspondence exists be-
process.
tween the set of coded characters
and the set of glyphs in a font.
7 Summary
–� The process is more complex
when it must choose between sev-
Here are the primary points of this technical
eral alternatives; for example,
report:
when a sequence of coded charac-
ters may be mapped into more
†� Most people equate a character and its
than one sequence of glyphs in a
shape.
font.
†� This causes difficulties and misunder-
standing because contemporary infor-
© ISO/IEC ISO/IEC TR 15285: 1998 (E)
Annex A
Bibliography
1. ISO/IEC 646: 1991, Information tech- 7. ANSI X3.4-1986, American National
nology — ISO 7-bit coded character set Standard for Information Systems —
for information interchange. Coded Character Sets — 7-Bit Ameri-
can National Standard Code for Infor-
2. ISO/IEC 6429: 1992, Information tech-
mation Interchange (7-Bit ASCII).
nology — Control functions for coded
character sets. 8. JIS X 0201-1976, Japanese Standards
Association, Jouhou koukan you fugou
3. ISO/IEC 6937: 1993, Information
(Code for Information Interchange).
technology — Coded graphic character
sets for text communication – Latin 9. JIS X 0208-1990, Japanese Standards
alphabet. Association, Jouhou koukan you kanji
fugoukei (Code of the Japanese
4. ISO/IEC 8859, Information technology
Graphic Character Set for Information
— 8-bit single-byte coded graphic
Interchange).
character sets
— Part 1. Latin alphabet No. 1 (1987) 10. Becker, Joseph D., “Multilingual Word
— Part 2. Latin alphabet No. 2 (1987) Processing”, Scientific American, Vol.
— Part 3. Latin alphabet No. 3 (1988) 251, No. 1, July, 1984, pp. 96–107.
— Part 4. Latin alphabet No. 4 (1988)
11. Bringhurst, Robert, The Elements of
— Part 5. Latin/Cyrillic alphabet (1988)
Typographic Style, Hartley and Marks,
— Part 6. Latin/Arabic alphabet (1987)
Vancouver, 1996.
— Part 7. Latin/Greek alphabet (1987)
— Part 8. Latin/Hebrew alphabet
12. Hartmann, R. R. K., and Stork, F. C.,
(1988)
Dictionary of language and linguistics,
— Part 9. Latin alphabet No. 5 (1989)
Applied Science Publishers Ltd., Lon-
— Part 10. Latin alphabet No. 6 (1993).
don, 1976.
5. ISO/IEC 10367: 1991, Information
13. Lofting, Peter, “The Perception of
technology — Standardized coded
Character Entities in Unfamiliar
graphic character sets for use in 8-bit
Scripts”, unpublished paper, July, 1995.
codes.
14. The Unicode Consortium, The Unicode
6. ISO/IEC 10538: 1991, Information
Standard, Version 2.0, Addison-
technology — Control functions for text
Wesley, Reading, MA, 1996.
communication.
ISO/IEC TR 15285: 1998 (E) © ISO/IEC
Annex B
Characters
posedly being represented by a character.
B.1 Definition
Instead, SC 2 assumes that the semantics
In ISO/IEC 10646-1:1993, SC 2 defines a of a character is either (1) self-evident or (2)
character as: subject to conventions adhered to by the
user of the character, namely, the
A member of a set of elements used
application.
for the organisation, control, and rep-
resentation of data. In a small character set standard, such as
ISO/IEC 646: 1991, the process of deter-
This definition asserts (1) that, in the con-
mining the information represented by each
text of the role of SC 2, a character is an
character is relatively straightforward and
element of a larger set, a character set, and
usually involves the invocation of self-
(2) that a character is used to represent
evident knowledge. For example, the char-
data or to organize and control data, or in a
acters of ISO/IEC 646 that appear to be the
few cases, both. The division between data
letters of the modern English alphabet, and
characters and control characters is usually
to which are assigned names that appear to
specified by requiring the former to be
be the names of the letters of this alphabet,
graphical characters, that is, characters with
are indeed usually assumed to represent
which some graphical form can be associ-
none other than the English alphabet. How-
ated. A character is not generally found (or
ever, this assumption is not supported by
interpreted) in isolation, but appears as an
the formal definition of ISO/IEC 646. No-
element of a sequence (an array) of charac-
where in this standard does it specify that
ters, that is, a character string, and there-
these characters actually represent informa-
fore is interpreted according to the context
tion to be interpreted as letters of the Eng-
in which it appears.
lish alphabet. Indeed, an application devel-
oper who happens to be Hawaiian may in-
After defining a character in this fashion, SC
terpret these characters as representing the
2 defines character sets by enumerating a
elements of the Hawaiian alphabet (plus a
list of characters. Such characters are enu-
few extra letters not used by Hawaiian), or a
merated by assigning a unique name to
Japanese developer may interpret them as
each character, by specifying a unique code
representing the elements of the Romaji
(the code position), and by depicting a rep-
form of written Japanese. In each case, the
resentative image in a table (the code ta-
user of the standard is applying conventions
ble). In general, this describes the entire
that do not conflict with the standard itself
formal content of any given SC 2 coded
and that enable the user to employ the
character set standard, although various
standard in a useful way. Other elements of
standards sometimes augment their formal
ISO/IEC 646, such as the character as-
content with additional information, particu-
signed to positions 2/13 (U+002D HYPHEN-
larly information pertaining to characters
MINUS “�”) and 2/7 (U+0027 APOSTROPHE “���
that participate in control functions.
”) are commonly given multiple interpreta-
tions depending on their use. For example,
B.2 Character information
the latter character may be used as an
apostrophe, as a single quote mark, or, in
What SC 2 does not do—and this is
some transliteration systems, as standing
perhaps the most important point of this
for a glottal stop or a palatalized consonant.
annex—is formally define the data or units
Since the standard does not specify which
of information that graphic characters are
information the character represents, a user
supposed to represent; that is, no formal
of the standard is free to choose. Once the
semantics are specified to assist in the task
number of characters in a standard is in-
of interpreting the so-called data sup-
creased many times, such as the case with
© ISO/IEC ISO/IEC TR 15285: 1998 (E)
ISO/IEC 10646-1: 1993 where over 30,000 Of these characters, the following are
characters are defined, the potential for merely size or position variants of a single
multiple usage conventions increases. form:
U+0031 DIGIT ONE “�”
B.3 Example, the unit of information
U+00B9 SUPERSCRIPT ONE “ ”
“one”
U+2081 SUBSCRIPT ONE “”
U+FF11 FULLWIDTH DIGIT ONE “�”
Consider for a moment the case of the unit
of information meaning “one”. ISO/IEC
The following are various adorned variants
10646 not only codes a large number of
of this form:
characters that conceivably represent this
unit of information but also codes a number
U+215F FRACTION NUMERATOR ONE “ ”
of characters that represent a particular U+2460 CIRCLED DIGIT ONE “�”
U+2474 PARENTHESIZED DIGIT ONE “���”
form associated with this meaning. The
U+2488 DIGIT ONE FULL STOP “��”
characters that may be said to represent the
U+2776 DINGBAT NEGATIVE CIRCLED DIGIT
unit of information designated by “one” are
ONE “�”
(at least):
U+2780 DINGBAT CIRCLED SANS-SERIF DIGIT
ONE “¥”
U+0031 DIGIT ONE “�”
U+278A DINGBAT NEGATIVE CIRCLED SANS-
U+00B9 SUPERSCRIPT ONE “”
SERIF DIGIT ONE “fl”
U+0661 ARABIC-INDIC DIGIT ONE “p”
U+06F1 EXTENDED ARABIC-INDIC DIGIT ONE
The following characters, although all rep-
“q”
resent the concept “one”, employ different
U+0967 DEVANAGARI DIGIT ONE “r”
forms depending on the script with which
U+09E7 BENGALI DIGIT ONE “s”
U+09F4 BENGALI CURRENCY NUMERATOR they are associated. However, one could
ONE “t”
argue that several of these forms are really
U+0A67 GURMUKHI DIGIT ONE “u”
different instances of a single form from
U+0AE7 GUJARATI DIGIT ONE “v”
which they are historically derived, namely,
U+0B67 ORIYA DIGIT ONE “w”
the Indic-script forms of “one”:
U+0BE7 TAMIL DIGIT ONE “x”
U+0C67 TELUGU DIGIT ONE “y”
U+0661 ARABIC-INDIC DIGIT ONE “p”
U+0CE7 KANNADA DIGIT ONE “z”
U+06F1 EXTENDED ARABIC-INDIC DIGIT ONE
U+0D67 MALAYALAM DIGIT ONE “{”
|“q”
U+0E51 THAI DIGIT ONE “ ”
U+0967 DEVANAGARI DIGIT ONE “r”
U+0ED1 LAO DIGIT ONE “}”
U+09E7 BENGALI DIGIT ONE “s”
U+2081 SUBSCRIPT ONE “”
U+0A67 GURMUKHI DIGIT ONE “u”
U+215F FRACTION NUMERATOR ONE “ ”
U+0AE7 GUJARATI DIGIT ONE “v”
U+2160 ROMAN NUMERAL ONE “,”
U+0B67 ORIYA DIGIT ONE “w”
U+2170 SMALL ROMAN NUMERAL ONE “L”
U+0BE7 TAMIL DIGIT ONE “x”
U+2460 CIRCLED DIGIT ONE “�”
U+0C67 TELUGU DIGIT ONE “y”
U+2474 PARENTHESIZED DIGIT ONE “���”
U+0CE7 KANNADA DIGIT ONE “z”
U+2488 DIGIT ONE FULL STOP “��”
U+0D67 MALAYALAM DIGIT ONE “{”
U+2776 DINGBAT NEGATIVE CIRCLED DIGIT
| U+0E51 THAI DIGIT ONE “ ”
ONE “�”
U+0ED1 LAO DIGIT ONE “}”
U+2780 DINGBAT CIRCLED SANS-SERIF DIGIT
U+3021 HANGZHOU NUMERAL ONE “~”
ONE “¥”
U+3192 IDEOGRAPHIC ANNOTATION ONE
U+278A DINGBAT NEGATIVE CIRCLED SANS-
MARK “a”
SERIF DIGIT ONE “fl”
U+3220 PARENTHESIZED IDEOGRAPH ONE
U+3021 HANGZHOU NUMERAL ONE “~”
“ ”
U+3192 IDEOGRAPHIC ANNOTATION ONE
U+3280 CIRCLED IDEOGRAPH ONE “ ”
MARK “a”
U+4E00 CJK UNIFIED IDEOGRAPH-4E00 “ ”
U+3220 PARENTHESIZED IDEOGRAPH ONE
U+58F9 CJK UNIFIED IDEOGRAPH-58F9 “ ”
“ ”
U+3280 CIRCLED IDEOGRAPH ONE “ ”
U+4E00 CJK UNIFIED IDEOGRAPH-4E00 “ ”
This example clearly shows that the de-
U+58F9 CJK UNIFIED IDEOGRAPH-58F9 “ ”
signers of this character set did not start
U+FF11 FULLWIDTH DIGIT ONE “�”
with individual units of information and as-
sign each such unit to a unique character;
ISO/IEC TR 15285: 1998 (E) © ISO/IEC
furthermore, it is also clear that the design-
ers did not start with individual forms and
assign each to a unique character. Rather,
a combination of forms and variations of a
Figure 6 — Old style figures
single form, all signifying the idea “one”,
were included as distinct characters.
B.4 Considerations for deciding the
repertoire of a coded character set
To gain an understanding of the distinction
between characters and glyphs, consider
Various arguments are possible for de-
that the following characters could have
fending the inclusion or exclusion of a par-
easily been unified into a single character
ticular form as a possible graphic character
that would be displayed using one of four
in a repertoire. In many cases, the criterion
glyphs:
for either inclusion or exclusion has not
been articulated but is based on informal
U+0031 DIGIT ONE “�”
opinion about appropriateness. Justifying
U+00B9 SUPERSCRIPT ONE “ ”
why certain forms were coded into ISO/IEC
U+2081 SUBSCRIPT ONE “”
10646-1: 1993 and why others were not is
U+FF11 FULLWIDTH DIGIT ONE “�”
beyond the scope of this Technical Report.
However, with respect to coding glyphs ver-
These four characters can be considered as
sus characters, the objective is to code
instances of one character that takes on
characters that represent different informa-
slightly different forms depending on usage.
tion. To meet this objective, three important
In this case, usage or style alone would
4)
considerations should be applied.
govern the form chosen to depict a single
abstract character. In the case of a form
1. Same shape/different meanings
used as the numerator of a fraction, the
appropriate glyph could be determined
Does one shape have multiple mean-
based on the local context of the character,
ings (semantics)?
assuming for a moment that a character
such as a U+0031 DIGIT ONE “�” is followed
Some shapes will be the same, or
by a U+2044 FRACTION SLASH “’”. In the
nearly the same, but have different
remaining cases, the character’s immediate
meanings or different semantics. An
context would not be sufficient but would
example of this is that in many sans-
require that additional information be sup-
serif fonts the glyph “I” is used for both
plied such a
...




Questions, Comments and Discussion
Ask us and Technical Secretary will try to provide an answer. You can facilitate discussion about the standard in here.
Loading comments...