ISO/IEC TR 24800-1:2007
(Main)Information technology — JPSearch — Part 1: System framework and components
Information technology — JPSearch — Part 1: System framework and components
ISO/IEC TR 24800-1:2007 specifies a framework for interoperability for still image search and retrieval and identifies an architecture and the components in this framework, the linkages between components, and which of these components and links are to be standardized in ISO/IEC 24800. ISO/IEC TR 24800-1:2007 specifies the following: Definitions of terms and abbreviations. A review of traditional approaches to image search and from various examples, motivates the importance of the user in the search process, the importance of making explicit the user task and user evaluation, and that meaning in images may be as much added from outside as extracted from the image itself. Real use cases of the various ways searching take place. In particular, there could be multiple entry points into the overall search framework. These include automatic, semi-automatic, and human (user) driven searches. A description of the overall search and management process. This can be considered the requirements specifications for a general search and management architecture. A 4-layer architecture for ISO/IEC 24800 and explicitly identifies the components in the architecture, and their roles and positions in the architecture. The overall structure of ISO/IEC 24800 and the purpose of the subsequent sub-parts: ISO/IEC 24800-2: Schema and ontology registration and identification ISO/IEC 24800-3: ISO/IEC 24800 query format ISO/IEC 24800-4: Metadata embedded in image data (JPEG-1 and JPEG-2000) file format ISO/IEC 24800-5: Data interchange format between image repositories
Technologies de l'information — JPSearch — Partie 1: Cadre système et composants
General Information
- Status
- Withdrawn
- Publication Date
- 13-Dec-2007
- Withdrawal Date
- 13-Dec-2007
- Current Stage
- 9599 - Withdrawal of International Standard
- Start Date
- 14-Aug-2012
- Completion Date
- 12-Feb-2026
Relations
- Effective Date
- 06-Aug-2011
Get Certified
Connect with accredited certification bodies for this standard

BSI Group
BSI (British Standards Institution) is the business standards company that helps organizations make excellence a habit.

NYCE
Mexican standards and certification body.
Sponsored listings
Frequently Asked Questions
ISO/IEC TR 24800-1:2007 is a technical report published by the International Organization for Standardization (ISO). Its full title is "Information technology — JPSearch — Part 1: System framework and components". This standard covers: ISO/IEC TR 24800-1:2007 specifies a framework for interoperability for still image search and retrieval and identifies an architecture and the components in this framework, the linkages between components, and which of these components and links are to be standardized in ISO/IEC 24800. ISO/IEC TR 24800-1:2007 specifies the following: Definitions of terms and abbreviations. A review of traditional approaches to image search and from various examples, motivates the importance of the user in the search process, the importance of making explicit the user task and user evaluation, and that meaning in images may be as much added from outside as extracted from the image itself. Real use cases of the various ways searching take place. In particular, there could be multiple entry points into the overall search framework. These include automatic, semi-automatic, and human (user) driven searches. A description of the overall search and management process. This can be considered the requirements specifications for a general search and management architecture. A 4-layer architecture for ISO/IEC 24800 and explicitly identifies the components in the architecture, and their roles and positions in the architecture. The overall structure of ISO/IEC 24800 and the purpose of the subsequent sub-parts: ISO/IEC 24800-2: Schema and ontology registration and identification ISO/IEC 24800-3: ISO/IEC 24800 query format ISO/IEC 24800-4: Metadata embedded in image data (JPEG-1 and JPEG-2000) file format ISO/IEC 24800-5: Data interchange format between image repositories
ISO/IEC TR 24800-1:2007 specifies a framework for interoperability for still image search and retrieval and identifies an architecture and the components in this framework, the linkages between components, and which of these components and links are to be standardized in ISO/IEC 24800. ISO/IEC TR 24800-1:2007 specifies the following: Definitions of terms and abbreviations. A review of traditional approaches to image search and from various examples, motivates the importance of the user in the search process, the importance of making explicit the user task and user evaluation, and that meaning in images may be as much added from outside as extracted from the image itself. Real use cases of the various ways searching take place. In particular, there could be multiple entry points into the overall search framework. These include automatic, semi-automatic, and human (user) driven searches. A description of the overall search and management process. This can be considered the requirements specifications for a general search and management architecture. A 4-layer architecture for ISO/IEC 24800 and explicitly identifies the components in the architecture, and their roles and positions in the architecture. The overall structure of ISO/IEC 24800 and the purpose of the subsequent sub-parts: ISO/IEC 24800-2: Schema and ontology registration and identification ISO/IEC 24800-3: ISO/IEC 24800 query format ISO/IEC 24800-4: Metadata embedded in image data (JPEG-1 and JPEG-2000) file format ISO/IEC 24800-5: Data interchange format between image repositories
ISO/IEC TR 24800-1:2007 is classified under the following ICS (International Classification for Standards) categories: 35.040 - Information coding; 35.040.40 - Coding of audio, video, multimedia and hypermedia information. The ICS classification helps identify the subject area and facilitates finding related standards.
ISO/IEC TR 24800-1:2007 has the following relationships with other standards: It is inter standard links to ISO/IEC TR 24800-1:2012. Understanding these relationships helps ensure you are using the most current and applicable version of the standard.
ISO/IEC TR 24800-1:2007 is available in PDF format for immediate download after purchase. The document can be added to your cart and obtained through the secure checkout process. Digital delivery ensures instant access to the complete standard document.
Standards Content (Sample)
TECHNICAL ISO/IEC
REPORT TR
24800-1
First edition
2007-12-15
Information technology — JPSearch —
Part 1:
System framework and components
Technologies de l'information — JPSearch —
Partie 1: Cadre système et composants
Reference number
©
ISO/IEC 2007
PDF disclaimer
This PDF file may contain embedded typefaces. In accordance with Adobe's licensing policy, this file may be printed or viewed but
shall not be edited unless the typefaces which are embedded are licensed to and installed on the computer performing the editing. In
downloading this file, parties accept therein the responsibility of not infringing Adobe's licensing policy. The ISO Central Secretariat
accepts no liability in this area.
Adobe is a trademark of Adobe Systems Incorporated.
Details of the software products used to create this PDF file can be found in the General Info relative to the file; the PDF-creation
parameters were optimized for printing. Every care has been taken to ensure that the file is suitable for use by ISO member bodies. In
the unlikely event that a problem relating to it is found, please inform the Central Secretariat at the address given below.
© ISO/IEC 2007
All rights reserved. Unless otherwise specified, no part of this publication may be reproduced or utilized in any form or by any means,
electronic or mechanical, including photocopying and microfilm, without permission in writing from either ISO at the address below or
ISO's member body in the country of the requester.
ISO copyright office
Case postale 56 • CH-1211 Geneva 20
Tel. + 41 22 749 01 11
Fax + 41 22 749 09 47
E-mail copyright@iso.org
Web www.iso.org
Published in Switzerland
ii © ISO/IEC 2007 – All rights reserved
Contents Page
Foreword. iv
Introduction . v
1 Scope .1
1.1 Interoperable Image Search and Retrieval.1
1.2 Motivation.1
1.3 Outline of the Technical Report .2
2 Terms, Definitions and Abbreviated Terms .3
2.1 Terms and definitions .3
2.2 Symbols and abbreviated terms .4
3 Background and motivation of a user-centric approach to image search .4
3.1 Traditional Models of Image Search and Retrieval .4
3.2 The user as part of the retrieval system.5
3.3 User task and user evaluation.6
3.4 Meaning also comes from outside.8
4 Use Cases.11
4.1 Introduction.11
4.2 Searching images in stock photo collections for usage in magazines .11
4.3 Searching for and publishing authoritative themed sub-collections of images.11
4.4 Mobile Tourist Information .11
4.5 Surveillance Search from Desktop to Mobile Device with Alerts .11
4.6 Ad hoc search without time-consuming housekeeping tasks.11
4.7 Rights clearance to publish a compliant business document.11
4.8 Tracking an object creation process using a temporal series of photos .11
4.9 Finding illegal or unauthorized use of images .11
4.10 Finding the best shots or filtering out of the worse shots.12
4.11 Context searching without human annotation .12
4.12 Image search based on image quality.12
4.13 Image search with deduplication .12
4.14 Matching images between collections for synchronization.12
4.15 Social metadata updating and sharing of images for searching.12
4.16 Image search in the medical domain.12
4.17 Servants image searchers .12
4.18 Open federated repositories.13
5 Image search and management process .13
5.1 Introduction.13
5.2 Metadata flow .13
5.3 Query process flow.15
6 JPSearch Architecture .18
6.1 Overview of the architecture .18
6.2 Parts to be standardized .20
6.3 Architecture justification with use cases.20
7 Organization of the JPSearch specification .21
7.1 Overall Structure of JPSearch.21
7.2 Part 2: Schema and ontology registration and identification .21
7.3 Part 3: JPSearch query format .22
7.4 Part 4: Metadata embedded in image data (JPEG and JPEG 2000) file format .22
7.5 Part 5: Data interchange format between image repositories .22
Annex A (informative) Use Cases .24
Bibliography .37
© ISO/IEC 2007 – All rights reserved iii
Foreword
ISO (the International Organization for Standardization) and IEC (the International Electrotechnical
Commission) form the specialized system for worldwide standardization. National bodies that are members of
ISO or IEC participate in the development of International Standards through technical committees
established by the respective organization to deal with particular fields of technical activity. ISO and IEC
technical committees collaborate in fields of mutual interest. Other international organizations, governmental
and non-governmental, in liaison with ISO and IEC, also take part in the work. In the field of information
technology, ISO and IEC have established a joint technical committee, ISO/IEC JTC 1.
International Standards are drafted in accordance with the rules given in the ISO/IEC Directives, Part 2.
The main task of the joint technical committee is to prepare International Standards. Draft International
Standards adopted by the joint technical committee are circulated to national bodies for voting. Publication as
an International Standard requires approval by at least 75 % of the national bodies casting a vote.
In exceptional circumstances, the joint technical committee may propose the publication of a Technical Report
of one of the following types:
— type 1, when the required support cannot be obtained for the publication of an International Standard,
despite repeated efforts;
— type 2, when the subject is still under technical development or where for any other reason there is the
future but not immediate possibility of an agreement on an International Standard;
— type 3, when the joint technical committee has collected data of a different kind from that which is
normally published as an International Standard (“state of the art”, for example).
Technical Reports of types 1 and 2 are subject to review within three years of publication, to decide whether
they can be transformed into International Standards. Technical Reports of type 3 do not necessarily have to
be reviewed until the data they provide are considered to be no longer valid or useful.
Attention is drawn to the possibility that some of the elements of this document may be the subject of patent
rights. ISO and IEC shall not be held responsible for identifying any or all such patent rights.
ISO/IEC TR 24800-1, which is a Technical Report of type [3], was prepared by Joint Technical Committee
ISO/IEC JTC 1, Information technology, Subcommittee SC 29, Coding of audio, picture, multimedia and
hypermedia information.
ISO/IEC TR 24800 consists of the following parts, under the general title Information technology — JPSearch:
⎯ Part 1: System framework and components
iv © ISO/IEC 2007 – All rights reserved
Introduction
JPSearch aims to provide a standard for interoperability for still image search and retrieval systems. There are
many systems that provide image search and retrieval functionality on computer desktops, on the World Wide
Web (i.e. websearch), on imaging devices, and in other consumer and professional applications. Existing
systems are implemented in a way that tightly couples many components of the search process. JPSearch
provides an abstract framework search architecture that decouples the components of image search and
provides a standard interface between these components.
Aligning image search system design to this standard framework facilitates the use and reuse of metadata;
the use and reuse of profiles and ontologies to provide a common context for searching; the provision of a
common query language to search easily across multiple repositories with the same search semantics; allows
image repositories to be independent of particular system implementations; and for users to move easily or
upgrade their image management applications or to move to a different device or upgrade to a new computer.
© ISO/IEC 2007 – All rights reserved v
TECHNICAL REPORT ISO/IEC TR 24800-1:2007(E)
Information technology — JPSearch —
Part 1:
System framework and components
1 Scope
1.1 Interoperable Image Search and Retrieval
This Technical Report specifies two things. The first is a framework for interoperability for still image search
and retrieval. The second identifies an architecture and the components in this framework, the linkages
between components, and which of these components and links are to be standardized in JPSearch.
The image search and retrieval framework will be determined by real use cases (tasks) and will leverage on
lessons learnt in the long history of text retrieval where, for example, different users issuing the same query
may be looking for (very) different results. This is important because it means that the framework must be
general enough to support many possible approaches to image retrieval, e.g., from using only low-level image
features, to text annotations, to community input, or a mixture of such approaches.
From the framework and components, and the linkages and flow of data between them, the parts of JPSearch
that need to be standardized can be determined.
1.2 Motivation
There are many applications that provide image search and retrieval functionality on computer desktops, on
the World Wide Web (i.e., websearch), on imaging devices, and in other consumer and professional
applications. These implementations are characterized by significant limitations, including:
• Lack of the ability to reuse metadata
The biggest problem in still image management is consistent and complete user or system annotation
(in whatever form) of images. A user makes a heavy investment if and when they annotate an image
or a collection of images. For example, a user adopts System A for storing and managing still images.
The user discovers System B, which provides improved and desirable functionality, but is effectively
prevented from switching to System B because the metadata in System A cannot be easily (or at all)
used in System B. In this example, users are impeded in using the applications or systems that best
suit their needs; and system providers are unable to compete freely with their products.
This problem generalizes in community based image sharing systems, where multiple users may
annotate the shared images. In most cases, however, an image has a single owner and there is a
need for the ability to merge community metadata back into the owner’s image management system.
This ability would help overcome the difficult problem of manual still image annotation.
• Lack of a common query format and search semantics
There is a trend towards shared image repositories. These could be on the web, but there are also
systems that publish user repositories residing on their local (e.g., home) machines for (normally
access controlled) public viewing and annotation. As the number and size of such repositories
increase (a monotonic increasing trend), search becomes an essential function for users to navigate
shared repositories.
© ISO/IEC 2007 – All rights reserved 1
Unfortunately, the various systems providing image search, whether on the desktop or on the
web, do not provide a common way of specifying a search. This is not the same as having a
common user-interface since the look and feel is up to a system provider to provide and for
the user to like or not like. The problem is that a query such as “white car” may be interpreted
as a Boolean “white” AND “car” or “white” OR “car”, or “white car” as a phrase, etc., and the
interpretation may be different when the search is done against the image data or against text
metadata, or against other metadata. Users are confused because different systems return
different results for the same query. System providers need a reference standard to remove
ambiguity and make searching over shared repositories consistent.
• Lack of a common format for handling context in searching
A large adult describing a 5-foot tall man may use the word “short”. A small child looking at
the same person may say “tall”. This does not mean that the person is both “short” and “tall”
at the same time; rather it is the context that has changed. Similarly, when a doctor does a
query using the term “skin cancer”, he or she probably expects a very different set of images
from when a patient searches with the same query term. Searching for images always takes
place in a context. This context may be implicit or explicit.
Some systems allow the user to specify a context and there are other systems that
automatically imply a user’s context. There is no way for the context in one search system to
be used in a different search system. A common format for handling context allows a user to
carry their context with them to different search engines. It also allows the context to be
owned by the user and not by the system, i.e., it protects the user’s privacy.
These are just three examples of where still image search systems can benefit tremendously from
interoperability. Other examples include how metadata can be created, evolved and stored, and also how
image collections can have metadata different from and augmenting the metadata of a single image.
Existing systems are implemented in a way that tightly couples many components of the search process.
JPSearch provides an abstract framework search architecture that allows an alignment of system design to a
standard framework. Among other things, this alignment facilitates the use and reuse of metadata, the use
and reuse of ontologies to provide a common language for contexts, the provision of a common query
language, provide standardized interface to system components, and the ability to provide still image search
and retrieval functionality across multiple repositories.
1.3 Outline of the Technical Report
There will be 7 clauses to this report. They are arranged as follows:
Clause 2 provides definitions of terms and abbreviations.
Clause 3 reviews the traditional approaches to image search and from various examples, motivates the
importance of the user in the search process, the importance of making explicit the user task and user
evaluation, and that meaning in images may be as much added from outside as extracted from the image
itself. This motivates the next clause, which are the real use cases.
Clause 4 describes the real use cases of the various ways searching does take place. In particular, there
could be multiple entry points into the overall search framework. These would include automatic, semi-
automatic, and human (user) driven searches. These would include specific use cases and motivating
examples of searches for images.
Clause 5 describes the overall search and management process. This can be considered the requirements
specifications for a general search and management architecture.
Clause 6 describes the 4-layer architecture for JPSearch and explicitly identifies the components in the
architecture, and what their roles and positions are in the architecture. We will describe how the use cases
described in Section 3 map to the layers of this architecture.
Clause 7 specifies the overall structure of JPSearch.
2 © ISO/IEC 2007 – All rights reserved
2 Terms, Definitions and Abbreviated Terms
For the purposes of this document, the following terms, definitions and abbreviated terms apply.
2.1 Terms and definitions
2.1.1
Annotation
metadata added to an image by way of definition or comment
NOTE It is normally in text and done by a human.
2.1.2
Content based retrieval system
system using non-text features of a content to search for similar contents
NOTE The abbreviated term CBIR is used to represent content-based retrieval system for image.
2.1.3
Communal recommendation systems
Systems which track items viewed by customers, and which groups customers by similar interests then
recommends items viewed by one customer to another customer in the same group
2.1.4
Context of a user
circumstances and conditions of the user during a query
2.1.5
Context of a query
context of the user as well as other circumstances and conditions affecting the query
2.1.6
Contextualization
process of placing a user or a process (e.g. a search) into a context
2.1.7
Index
way of organizing data that improves searching the data
EXAMPLE A library catalogue is an index. It is normally sorted alphabetically to make it more efficient to find an
entry in the index. Indexing is the process of creating an index.
2.1.8
Metadata
data about data
EXAMPLE An image is a data item. Metadata about the image may include information such as the size of the
image, the date it was created, etc.
2.1.9
Ontology
model that represents a domain and is used to reason about the objects in that domain and the relations
between them
NOTE This is a form of knowledge representation about the world or some part of it.
2.1.10
Pragmatics
leftover part in a theory of language and communication after the syntax and semantics have been taken out
© ISO/IEC 2007 – All rights reserved 3
2.1.11
Query
request for information from a search and retrieval system
2.1.12
Query expansion
technique in information retrieval that adds terms to a query to improve the accuracy of the search or to
increase the number of useful documents retrieved
2.1.13
Query-by-example
type of query where an example of the answer desired is used as the input to the search system
2.1.14
Reverse Index
type of index where, given a word, one can look up the reverse index and locate all the documents where the
word occurs, and optionally, where in the document and how often the word occurs
2.1.15
Semantics
mapping between elements of a language and the real world
2.1.16
Syntax
set of rules that govern whether a sentence (or other unit of communication) is well formed
2.2 Symbols and abbreviated terms
ALT: Alternative Text
CBIR: Content-Based Image Retrieval
HTML: Hypertext Markup Language
QBE: Query-by-example
UI: User Interface
USB: Universal Serial Bus
3 Background and motivation of a user-centric approach to image search
3.1 Traditional Models of Image Search and Retrieval
To put things into perspective, we need to start with the field of document or text search and retrieval. This
has a long rich history rooted in the field of library science and has evolved through forms of text processing,
information tracking, through to document “retrieval engines”, leading to the various web search engines.
Figure 1(a) shows the naive system view for document search and retrieval that made up the early retrieval
systems. Documents in the collection to be searched had first to be indexed. This was commonly to treat them
as a “bag of words” that were then efficiently stored in a reverse index. A query is normally a bunch of
keywords that is matched against the index and the appropriate documents retrieved.
Figures 1(b) and 1(c) show the traditional approaches to digital image searching. This has primarily fallen into
two camps. The first, as shown in Figure 1(b), is to search by keywords, i.e., it requires each image to be
associated with one or more keywords. There are various schemes to associate the keywords with the images
but the most effective has been manual annotation, e.g., done by home users for their home photos or by
trained domain experts in more commercial settings. As can be seen from comparing 1(a) and 1(b), it is
essentially the same kind of naive document search system.
4 © ISO/IEC 2007 – All rights reserved
Figure 1 — Naive system view of information retrieval, Image search, and CBIR.
The second common form of image retrieval uses an image as a query and the system attempts to retrieve
other images that are similar (Figure 1(c)). This is the accepted state of the art in content-based information
retrieval (CBIR) systems. The primary area of research is oriented around discovering and extracting new
kinds of image features to characterize the image for better performance during retrieval. Some of these
features include color (in the form of color histograms, color moments, color deltas, etc.), edges (whether as
line features or “assembled” into higher-order objects), texture, blobs, regions, etc. While the features have
become more sophisticated, fundamentally the systems follow the model in Figure 1(c). Again as can be seen
by comparing with Figures 1(a) and 1(b), there is little difference with the other naive systems.
There are also hybrid systems that combine text and image queries for searching. Text queries are used as
an entry point into a search or browse space, after which image-to-image matching is used to refine the query
or to retrieve further “similar” results. The system as a whole becomes more complex, but each of the
components still behaves as above.
3.2 The user as part of the retrieval system
Figure 1 showed the similarity of current image retrieval models to the naive models of document retrieval
prevalent about 20 years ago. Since that time, document and text retrieval has improved by leaps and
bounds. One of the clear factors in this evolution is the recognition that the user should be treated as an
integral part of the information retrieval process, i.e., that finding the right information is about much more than
just the search system provided by a library or by a vendor or on a website.
Figure 2 — Adding the user into the system model
© ISO/IEC 2007 – All rights reserved 5
Having the user in the loop provides many advantages to a search system. One of the biggest problems in
search is ambiguity i.e., which meaning of a word, or which aspect of an image, is the one that is appropriate
in a given search situation. Consider a chair. If we were to sit on it, it would be a chair. If we wanted to buy
one, we couldn’t go to a “chair store”; we have to go to a furniture store. And assuming it’s a wooden chair, if
we were stuck in a blizzard and feeling really cold, it would be firewood! Different aspects of an object, or a
picture, or a document are appropriate in different circumstances. These circumstances could depend on the
user’s needs (I want to buy one), the context (I’m really cold), the utility (firewood), or any one or more of a
myriad other conditions. The converse of ambiguity is redundancy where one thing can be referred to by many
names, e.g., the planet Venus was also known as the Morning Star and as the Evening Star, or your portable
computer being a laptop or else a notebook computer.
Having the user as part of the system gives us the model shown in Figure 2. The indexing side hasn’t
changed; the query process is a little more complicated. The user has an initial query (step 1) for which the
system returns a set of results, just as in the naive system. However, here the evaluation of the usefulness of
the results is done by the user (step 2) taking into account his or her own circumstances, etc. If necessary, the
query can be modified (step 3) and the search done again.
The user also has other influences on how the search system operates. One important recognition is that
there is not just one kind of search; that users engage in different Information Seeking Strategies (ISS) [Belkin
et.al, 1995]. Four dimensions of ISS were identified, specifically Method of Interaction, Goal of Interaction,
Mode of Retrieval, and Resource Considered. To quote,
“method of interaction, can be understood in terms of the classic distinction
between searching for a known item and looking around, or scanning, for
something interesting among a collection of items. The goal of the interaction
may be learning about some aspect of an item or resource, or selecting
useful items for retrieval. Furthermore, looking for identified items can be
characterized as retrieval by specification, while identifying relevant items
through stimulated association can be characterized as retrieval by
recognition. And interaction with information items themselves can be
contrasted with interaction with meta-information resources that describe the
structure and contents of information objects”. [Belkin et.al, 1995, pg 385]
Each of these dimensions represents a choice by the user as they use the search system. If the system does
not support that choice, then the user is, in a sense, fighting with the system to achieve his or her goals. For
example, traditional libraries (the ones with books on shelves) easily support scanning (and consequently
serendipity) whereas digital libraries (with only e-books) have to come up with alternatives (e.g., communal
recommendation systems).
3.3 User task and user evaluation
In the previous section, we have motivated the various reasons for including the user in the search process.
This section looks at the reason for a user to go to a search system in the first place. One does not just decide
for no reason to go to a search engine and look for something. A user must start with a purpose, something he
or she wants done, and in the course of doing it or planning on how to do it, knows (or finds out) that there is
missing information that interferes with the successful completion of what is to be done. This has been
characterized, for example, as a state of anomalous knowledge [Belkin, et. al, 1982]. It is important to note
that the user is in a state, for this implies a larger context within which one or more pieces of information are
missing. The user goes to the search system to find those missing pieces of information. Finding the missing
information and thus allowing the purpose to be satisfied is called the task.
6 © ISO/IEC 2007 – All rights reserved
Figure 3 — Results from a typical search engine for Eiffel tower restaurant
A typical search engine returns a list of results. These are often ranked, i.e., the topmost result is what the
engine predicts to be the most relevant to the given query. Note that this is with respect to a query and not to
a user or a task. The change in paradigm that comes with making the user a part of the information retrieval
system is that relevance is now defined with respect to the user and the task. This is called user evaluation.
Consider the two following scenarios. John is attending a convention in Las Vegas and wants to find the
phone number of the Eiffel Tower Restaurant at the Paris Hotel to make a dinner reservation. Betty is a
student trying to find the name of the restaurant on the Eiffel Tower (the one in Paris, France) observation
deck for a school project. Here we see two different users each with a very distinct task. But when it comes to
the search engine, they enter identical query terms. Both unsurprisingly type in “Eiffel tower restaurant”. The
first 6 results of the search are shown in Figure 3.
If you take the naive system view, i.e., that you evaluate relevance with respect to the query, then all 6 results
returned were “relevant” since they were all about Eiffel tower restaurant in some way or another. However,
when you include the user and the task in the evaluation, this changes very dramatically. With user evaluation,
st th th
the 1 , 5 and 6 results are relevant to John, and the rest are relevant to Betty, i.e., only 50% of the
documents returned are relevant in either case, and none of the documents relevant to John in his task was
th rd
relevant to Betty in her task. Incidentally, notice that the 6 and 3 results respectively already satisfy John’s
and Betty’s tasks even without having to open the documents (websites).
© ISO/IEC 2007 – All rights reserved 7
Figure 4 — When you fix the user and the task, the model reduces to that in Figure 1
A text example was used above for clarification, but the idea of user task and user evaluation applies equally
well for digital image searching. The state of the art in image searching research does include a user and a
task. However, to make things less subjective, researchers hold the user and the task as constants. For a
given query, they are then able to decide which images are “relevant” and which are not. For our example
above, this would be to say that either the answers for John are “right” xor those for Betty are “right”, without
considering (and supporting) the possibility that both sets of answers could be “right” in the appropriate
circumstances. In image retrieval systems, this fixed user and fixed task evaluation is called the ground truth,
and is used as an objective standard to measure the performance of an image search engine. While this
makes the comparison between different search systems easier, once you fix the user and task, it gives rise
to the situation in Figure 4, i.e., the system reduces back to the model in Figure 1 (the naive system view).
What this means is that the system measures relatively well for a given user and a given task, but only for that
given user in that given task. A generic model for digital image searching must be able to handle different
users (with their respective problems they want to solve) and different, even conflicting, tasks at the same
time.
3.4 Meaning also comes from outside
Semantics is a misnomer in the way it is used in the image (or video) searching community. Linguistically
speaking [Newmeyer1986], semantics refers to the mapping between well-formed expressions in a language
and the real world. It is invariably compositional, i.e., if I know what a and b denote and I know what the
operator + denotes, then I know what a+b denotes. However this doesn’t work for image “semantics”. Just
because I know that this region is a blob, that it’s round and that it’s red doesn’t mean I know that it’s a
tomato. Or a red ball. Or the sun. Or whatever else could be round and red. In fact, the sun has often been
referred to as a red ball, so these are not exclusive possibilities. Properly, the idea of trying to make sense out
of an image is closer to what linguists term pragmatics rather than semantics. Pragmatics is very much what it
seems to be [Levinson 1983]. Without worrying about formal definitions of atomic units (pixels? blobs?) or
compositional rules (2 adjacent pixels of the same color = a blob of that color?), pragmatics tries to figure out
how two people are able to communicate even if they are using (or misusing) a language with nonsense
referents and non-well-formed utterances.
8 © ISO/IEC 2007 – All rights reserved
Figure 5 — picture taken from www.whitehouse.gov
The distinction between semantics and pragmatics is important. Linguistically, semantics is defined with
respect to the utterance only, i.e., it doesn’t matter what the speaker of the utterance meant when he made
the utterance; nor does it depend on what the listener took away from hearing the utterance. There is, in other
words, the same idea of a user-independent ground truth. Pragmatics, on the other hand, is very dependent
on the speaker and the listener. Consider Joan and Darby, who are a long married couple and they’re at the
Cineplex. They’re standing looking at the list of movies showing and after a couple of minutes; Darby turns to
Joan and says, “So?” That’s not a well-formed sentence and has no semantics, but that single utterance,
pragmatically, carried a wealth of meaning. Darby, by that one-word utterance, was saying to Joan, “there’s
your favorite romantic weepy showing that always puts me to sleep, and the usual blockbuster action movie
that gives you a headache, so are we going to settle for our normal campy comedy or try out that animated
children’s film?”. This is not an unbelievable scenario by any means. What it shows is that once you have the
human user in the loop, it’s pragmatics and not semantics that is important, and that pragmatics needs the
intelligence and the interpretation of the human to make it work. Where there isn’t many years of marriage to
provide a shared context, there is often redundancy, social norms, and shared objectives that fulfil the same
role.
So, this brings us to image information retrieval. In text retrieval, a document is effectively indexed as a bag of
words, without reference to outside pragmatics. There are heuristics that are useful, e.g., that the more times
a word occur in a document, the more likely it is that the word is a better proxy for the document during a
search. In text-based image retrieval (i.e., using annotations), the annotated text is the proxy, and once the
annotations are entered, there is again no reference to outside pragmatics. Where the images occur in
context, e.g., on a webpage, then the caption, the HTML ALT text for the image, or the text on the page can
be and are used as proxies. However, there may be smart schemes to create “better” annotations, or to
automatically create annotations for the text. These then require information to be added to the image.
Consider Figure 5. If the annotation was “man using a telephone”, it might be argued that this is a very smart
extraction of semantics from the image. An even smarter one would be “George Bush in the Oval Office
talking on the phone”. The actual caption (annotation) of the picture says “President George W. Bush speaks
via phone to Associate Supreme Court Justice Sandra Day O'Connor Friday, July 1, 2005, shortly after she
submitted her letter of resignation citing personal reasons. The letter sits on the desk.” That annotation
contains information that obviously cannot be extracted from the image, including (a) that the date is 1 July,
(b) that the person on the other end of the phone is Sandra Day O’Connor, and (c) that she had just resigned,
etc.
© ISO/IEC 2007 – All rights reserved 9
Figure 6 — Microscope slide of oral tissue showing cancer growth
Lastly in content-based image retrieval (CBIR), we have multiple layers of meaning all potentially active at the
same time. There are primary features that can be extracted from an image. These are unambiguous
information that is structurally (or syntactically) objective such as color histogram, texture, lines, etc. This has
also been called low-level features in the CBIR literature. There are also secondary features that can be
extracted from an image, especially in specialized domains such as medical image analysis [Muller et.al
2004]. These are features which form is intrinsic to the image, but which significance comes from outside. For
example, a particular texture from a tissue image on a medical slide may have a high correlation with cancer
(see Figure 6).
Then there are higher level of “semantics” that come from analysis and deduction (e.g., face recognition in
images). Looking at Figure 5 again, if a face recognition tool identified the person in the image as “George
Bush”, it would require adding information from outside the image to the meaning of the image itself. Consider
the following text taken from a news source:
"The president of the United States said in South Korea that the United
States has no intention to attack North Korea. They’ve been told they can
have multilateral security assurances if they will make the important decision
to give up their nuclear weapons program"
If after reading the above statement, I were to summarize the text by saying “George Bush says that the US
will not attack N. Korea”, I would also have had to add information in that summary, specifically that George
Bush denotes the same person (at this place and time) as the definite description, the president of the United
States.
Another kind of higher level “semantics” comes from looking at similarities that occur across images, i.e.,
within a class of images [Lim & Jin, 2004a, 2004b]. So pictures of beach scenes have similar features
because they have similar content. One blue blob in an image may not have any useful meaning. But when
similar blue blobs occur in many pictures, then it becomes interesting. We can sometimes give labels to these
types of natural classes, e.g., for the class of beach scenes, we have sky, sea, cloud, sand, etc., though there
are sometimes patterns in such image collections that the system may discover but which may not correspond
to human labels.
In summary, we have shown that there are two ways of ascribing meaning to an image. The simple one is to
consider only the intrinsic image content. The more sophisticated version uses a lot of external knowledge,
context, or reasoning. In particular, these sophisticated methods may require considering the image content in
comparison with other images in a class, or with respect to extrinsic prior knowledge. In these and other
cases, meaning is added to the image rather than being extracted from an image.
10 © ISO/IEC 2007 – All rights reserved
4 Use Cases
4.1 Introduction
From the previous clause, we see that the JPSearch framework architecture needs to be able to handle very
diverse ways of ascribing semantics (i.e., meaning in the form of metadata, context, profiles, etc.) to images,
and also to collections of images. Thus many use cases were solicited from consumer, professional and
academic sources. These were discussed and analyzed and the following use cases were deemed to be
representative of the wide spectrum of user needs as discussed in clause 4.
The following descriptions of use cases are only summaries. The full description of each use case is provided
in Annex A.
4.2 Searching images in stock photo collections for usage in magazines
The user wishes to buy a selection of images in order to illustrate a publication to be sold to consumers.
4.3 Searching for and publishing authoritative themed sub-collections of images
The user has specific directives and procedures to annotate and manage themed collections of digital images,
e.g., “Fake Boticelli Paintings”. The collections will then be accessed by third parties (e.g., professional users
purchasing material for commercial use). In particular, this refers to creating sub-collections of themes that are
not explicitly annotated in the repository.
4.4 Mobile Tourist Information
A user (tourist) is in an unfamiliar place, sees an interesting landmark and wants to know what it is. He takes a
picture of the landmark on his mobile phone and sends it to a tourist information server that calls him back and
gives him the information.
4.5 Surveillance Search from Desktop to Mobile Device with Alerts
The user sets up a visual surveillance query on a desktop computer (large screen, comfortable keyboard),
saves the query for real time monitoring with results saved periodically for retrieval from a mobile device. In
addition, a trigger event occurs, a short message alert is sent to the user.
4.6 Ad hoc search without time-consuming housekeeping tasks
Users wish to bring their private photo collection on their personal storage devices such as a memory card
and to retrieve images using any ter
...




Questions, Comments and Discussion
Ask us and Technical Secretary will try to provide an answer. You can facilitate discussion about the standard in here.
Loading comments...