ETSI TS 104 286 V1.1.1 (2026-08)
Methods for Testing & Specification (MTS); Methodology Definition; Digital Transformation of Security Standards into Requirements
General Information
- Abstract
DTS/MTS-TST12TraStaReq
- Status
- Not Published
- Technical Committee
- MTS TST - Testing
- Current Stage
- 12 - Citation in the OJ (auto-insert)
- Due Date
- 04-Sep-2026
- Completion Date
- 07-Aug-2026
Frequently Asked Questions
ETSI TS 104 286 V1.1.1 (2026-08) is a standard published by the European Telecommunications Standards Institute (ETSI). Its full title is "Methods for Testing & Specification (MTS); Methodology Definition; Digital Transformation of Security Standards into Requirements". This standard covers: DTS/MTS-TST12TraStaReq
DTS/MTS-TST12TraStaReq
ETSI TS 104 286 V1.1.1 (2026-08) is available in PDF format for immediate download after purchase. The document can be added to your cart and obtained through the secure checkout process. Digital delivery ensures instant access to the complete standard document.
Standards Content (Sample)
TECHNICAL SPECIFICATION
Methods for Testing & Specification (MTS);
Methodology Definition;
Digital Transformation of Security Standards into
Requirements
2 ETSI TS 104 286 V1.1.1 (2026-08)
Reference
DTS/MTS-TST12TraStaReq
Keywords
IoT, LLM, methodology, security
ETSI
650 Route des Lucioles
F-06921 Sophia Antipolis Cedex - FRANCE
Tel.: +33 4 92 94 42 00 Fax: +33 4 93 65 47 16
Siret N° 348 623 562 00017 - APE 7112B
Association à but non lucratif enregistrée à la
Sous-Préfecture de Grasse (06) N° w061004871
Important notice
The present document can be downloaded from the
ETSI Search & Browse Standards application.
The present document may be made available in electronic versions and/or in print. The content of any electronic and/or
print versions of the present document shall not be modified without the prior written authorization of ETSI. In case of any
existing or perceived difference in contents between such versions and/or in print, the prevailing version of an ETSI
deliverable is the one made publicly available in PDF format on ETSI deliver repository.
Users should be aware that the present document may be revised or have its status changed,
this information is available in the Milestones listing.
If you find errors in the present document, please send your comments to
the relevant service listed under Committee Support Staff.
If you find a security vulnerability in the present document, please report it through our
Coordinated Vulnerability Disclosure (CVD) program.
Notice of disclaimer & limitation of liability
The information provided in the present deliverable is directed solely to professionals who have the appropriate degree of
experience to understand and interpret its content in accordance with generally accepted engineering or
other professional standard and applicable regulations.
No recommendation as to products and services or vendors is made or should be implied.
No representation or warranty is made that this deliverable is technically accurate or sufficient or conforms to any law
and/or governmental rule and/or regulation and further, no representation or warranty is made of merchantability or fitness
for any particular purpose or against infringement of intellectual property rights.
In no event shall ETSI be held liable for loss of profits or any other incidental or consequential damages.
Any software contained in this deliverable is provided "AS IS" with no warranties, express or implied, including but not
limited to, the warranties of merchantability, fitness for a particular purpose and non-infringement of intellectual property
rights and ETSI shall not be held liable in any event for any damages whatsoever (including, without limitation, damages
for loss of profits, business interruption, loss of information, or any other pecuniary loss) arising out of or related to the use
of or inability to use the software.
Copyright Notification
No part of this document may be reproduced in any form, by any means and in any media, without the prior written
authorization of ETSI and except as expressly permitted below.
By way of exception and when the document is a normative deliverable (European Standard (EN),
Technical Specification (TS), Group Specification (GS) or ETSI Standard (ES)), ETSI authorizes to reproduce
and incorporate into products, services and technical documentation only those extracts (e.g. templates) that are strictly
necessary for the technical implementation of the normative deliverable, to ensure compliance with the latter.
Nothing in this notice shall be construed as limiting any mandatory exceptions to copyright provided by applicable law.
© ETSI 2026.
All rights reserved.
ETSI
3 ETSI TS 104 286 V1.1.1 (2026-08)
Contents
Intellectual Property Rights . 4
Foreword . 4
Modal verbs terminology . 4
Executive summary . 4
Introduction . 5
1 Scope . 7
2 References . 7
2.1 Normative references . 7
2.2 Informative references . 7
3 Definition of terms, symbols and abbreviations . 8
3.1 Terms . 8
3.2 Symbols . 8
3.3 Abbreviations . 9
4 Background . 9
5 Methodology Overview . 9
6 Methodology . 10
6.1 Workflow . 10
6.2 Adaptation Techniques . 11
6.3 Document Parsing (DP) . 12
6.4 Requirement Identification and Extraction (RIE) . 12
6.4.1 Fine-tuning . 12
6.4.1.1 Overview . 12
6.4.1.2 Dataset Construction and Data Preparation . 13
6.4.1.3 Model Selection . 14
6.4.1.4 Model Evaluation . 14
6.4.1.5 Model Execution . 14
6.4.2 Prompt Engineering . 15
6.5 Requirement Classification (RC) . 15
6.5.1 Fine-tuning . 15
6.5.1.1 Overview . 15
6.5.1.2 Dataset Construction and Data Preparation . 15
6.5.1.3 Preprocessing and Model Construction . 16
6.5.1.4 Model Evaluation . 16
6.5.1.5 Model Execution . 16
6.5.2 Prompt Engineering . 16
6.6 Requirement Checklist Generation (RCG) . 17
6.7 Feedback Provision (FP) . 17
Annex A (normative): Examples of extracted requirements . 19
History . 23
ETSI
4 ETSI TS 104 286 V1.1.1 (2026-08)
Intellectual Property Rights
Essential patents
IPRs essential or potentially essential to normative deliverables (European Standard (EN), Technical Specification (TS),
Group Specification (GS) or ETSI Standard (ES)) may have been declared to ETSI. The declarations pertaining to these
essential IPRs, if any, are publicly available for ETSI members and non-members, and can be found in
ETSI SR 000 314: "Intellectual Property Rights (IPRs); Essential, or potentially Essential, IPRs notified to ETSI in
respect of ETSI standards", which is available from the ETSI Secretariat. Latest updates are available on the
ETSI IPR online database.
Pursuant to the ETSI Directives including the ETSI IPR Policy, no investigation regarding the essentiality of IPRs,
including IPR searches, has been carried out by ETSI. No guarantee can be given as to the existence of other IPRs not
referenced in ETSI SR 000 314 (or the updates on the ETSI Web server) which are, or may be, or may become,
essential to the present document.
Trademarks
The present document may include trademarks and/or tradenames which are asserted and/or registered by their owners.
ETSI claims no ownership of these except for any which are indicated as being the property of ETSI, and conveys no
right to use or reproduce any trademark and/or tradename. Mention of those trademarks in the present document does
not constitute an endorsement by ETSI of products, services or organizations associated with those trademarks.
DECT™, PLUGTESTS™, UMTS™ and the ETSI logo are trademarks of ETSI registered for the benefit of its
Members. 3GPP™, LTE™ and 5G™ logo are trademarks of ETSI registered for the benefit of its Members and of the
3GPP Organizational Partners. oneM2M™ logo is a trademark of ETSI registered for the benefit of its Members and of ®
the oneM2M Partners. GSM and the GSM logo are trademarks registered and owned by the GSM Association.
Foreword
This Technical Specification (TS) has been produced by ETSI Technical Committee Methods for Testing and
Specification (MTS).
Modal verbs terminology
In the present document "shall", "shall not", "should", "should not", "may", "need not", "will", "will not", "can" and
"cannot" are to be interpreted as described in clause 3.2 of the ETSI Drafting Rules (Verbal forms for the expression of
provisions).
"must" and "must not" are NOT allowed in ETSI deliverables except when used in direct citation.
Executive summary
The present document introduces a methodology for the semi-automated transformation of security standards into a list
of security requirements. More specifically, it describes an AI-driven methodology for the identification, extraction, and
classification of security requirements from text retrieved from security standards documents. The methodology is
based on adapting pre-trained foundational Large Language Models (LLMs) using common adaptation techniques like
prompt engineering and fine-tuning in the downstream tasks of security requirements identification, extraction, and
classification. The present document can also be used as a guide for practitioners who are interested in utilizing LLMs
for automating the identification, extraction, and/or classification of security requirements from security standards.
ETSI
5 ETSI TS 104 286 V1.1.1 (2026-08)
Introduction
Compliance with international security standards is essential for ensuring the security of digital products and services,
and thereby their dependability and trustworthiness. Beyond fostering trust, compliance with security requirements is
increasingly mandated by regulatory frameworks in the European Union (EU), such as the Cyber Resilience Act
(CRA) [i.1]. In particular, under the CRA, products with digital elements should comply with critical product-level
cybersecurity requirements before being placed on the EU market. The CRA entered into force on 10 December 2024,
and its main obligations apply from 11 December 2027 (with certain obligations applying earlier). Therefore,
compliance with international security standards and regulations has become highly relevant to a broad set of
stakeholders and application domains, including some that previously had fewer formal cybersecurity obligations.
Compliance evaluation is typically conducted by security experts (e.g. internal assessors, auditors, etc.) who verify
whether critical security requirements expressed in relevant international security standards or regulations are satisfied
by the subject system. To state it simply, the security expert works from a list of security requirements that the system
needs to satisfy and gathers objective evidence (e.g. technical documentation, test results, etc.) regarding the satisfaction
of these requirements. If sufficient evidence is found for the considered requirements, the system can be regarded as
fully compliant with the designated/selected security standard or regulation.
Compliance evaluation starts with the extraction of the security requirements that are expressed in a relevant standard or
regulation that the system needs to comply with, and their representation in the form of a list of clearly-described and
well-structured requirements. Extracting requirements from standards is a tedious, time-consuming, and effort-
demanding task, which is also prone to human error. This is mainly attributed to the fact that international security
standards are often long documents with complex, highly technical language, requiring technical knowledge and
expertise by the person who is tasked to go through their text and detect and extract the security requirements. Until
today, this process remains largely manual. Although several rule-based and Natural Language Processing (NLP)
techniques have been introduced, they have not achieved the accuracy and generalizability required for broader
adoption in practice.
However, the recent advancements in the field of Artificial Intelligence (AI) and particularly the proposition of the
Transformer architecture and its derivative LLMs that have demonstrated remarkable capabilities in natural language
understanding and processing open a new potential for the automation of the digital transformation process of security
standards and regulations.
The present document introduces an AI-based methodology for the semi-automated digital transformation of security
standard documents into lists of actionable security requirements, which can be used as the foundation for conducting
subsequent compliance evaluation activities. The methodology is based on adapting LLMs in order to be able to:
i) identify and extract security requirements that are described in the documents of international security
standards and regulations;
ii) categorize/classify them into high-level security criteria/characteristics to which they belong (e.g.
Confidentiality, Integrity, Availability, etc.); and
iii) extract the identified requirements in a unified form, which is both machine- and human-readable.
The methodology for the digital transformation of security standards is based on two separate LLM-based mechanisms,
one for the identification and extraction of the security requirements from text retrieved from international security
standards and regulations, and another for the classification of the extracted requirements into specific security
categories. These two mechanisms are used jointly in order to deliver the desired functionality, acting as independent
collaborative AI agents, each one excelling in its designated task and acting synergistically in order to achieve the high-
level goal of turning a standard document into a well-structured list (checklist) of security requirements.
The proposed methodology is based on the adaptation of pre-trained foundational LLMs, leveraging common
adaptation techniques, including prompt engineering (both zero- and few-shot learning) and fine-tuning. The most
suitable adaptation technique needs to be selected for each one of the considered tasks, i.e. requirements identification,
extraction, and classification. The present document, apart from the methodology itself, provides guidelines on how to
adapt existing LLMs in order to be used as the basis for these important requirements engineering tasks.
The methodology is intended for use by conformity assessment bodies and auditors preparing requirement checklists to
be used as the basis for subsequent compliance and certification activities, by product and security engineering teams
deriving security requirements from applicable standards, by standards owners and working groups piloting semi-
automated requirement-extraction workflows, by tool developers and integrators building LLM-assisted compliance
tooling, and by regulators and market-surveillance authorities exploring repeatable methods for requirements derivation.
ETSI
6 ETSI TS 104 286 V1.1.1 (2026-08)
The present document can be used for deriving a clearly structured set of requirement records from a referenced
standard or regulation, configuring and adapting LLMs for requirement identification, extraction and classification,
establishing traceability from each derived requirement to the source clause for auditability, supporting
human-in-the-loop review and quality control, harmonising internal compliance datasets and processes across products
and projects, and benchmarking alternative prompts, models and settings against common acceptance criteria.
The present document is independent of any specific product, vendor, model family, deployment environment or
programming framework. It does not prescribe a particular LLM, model size, decoding strategy, training corpus, cloud
provider or tooling stack, and it is compatible with both proprietary and open-source implementations, whether
deployed on-premises or in cloud environments. Any examples or parameter values are illustrative only and do not
imply endorsement. References to external documents are non-exhaustive and provided for context.
ETSI
7 ETSI TS 104 286 V1.1.1 (2026-08)
1 Scope
The present document introduces an AI-driven methodology for enabling the semi-automated transformation of
international security standards into a comprehensive list of security requirements. More specifically, the present
document specifies a methodology for:
a) The automatic identification and extraction of security requirements from text retrieved from standard
documents.
b) The classification of security requirements extracted from standards into critical high-level security
categories/attributes (e.g. Confidentiality, Integrity, Availability, etc.).
c) The semi-automated transformation of complete standard documents into a list of well-structured security
requirements, incorporating a feedback loop with the user.
d) Adapting LLMs in the tasks of security requirements identification, extraction and classification.
2 References
2.1 Normative references
References are either specific (identified by date of publication and/or edition number or version number) or
non-specific. For specific references, only the cited version applies. For non-specific references, the latest version of the
referenced document (including any amendments) applies.
Referenced documents which are not found to be publicly available in the expected location might be found in the
ETSI docbox.
NOTE: While any hyperlinks included in this clause were valid at the time of publication, ETSI cannot guarantee
their long-term validity.
The following referenced documents are necessary for the application of the present document.
Not applicable.
2.2 Informative references
References are either specific (identified by date of publication and/or edition number or version number) or
non-specific. For specific references, only the cited version applies. For non-specific references, the latest version of the
referenced document (including any amendments) applies.
NOTE: While any hyperlinks included in this clause were valid at the time of publication, ETSI cannot guarantee
their long-term validity.
The following referenced documents may be useful in implementing an ETSI deliverable or add to the reader's
understanding, but are not required for conformance to the present document.
[i.1] Regulation (EU) 2024/2847 of the European Parliament and of the Council of 23 October 2024 on
horizontal cybersecurity requirements for products with digital elements and amending
Regulations (EU) No 168/2013 and (EU) 2019/1020 and Directive (EU) 2020/1828 (Cyber
Resilience Act), OJ L, 20.11.2024.
[i.2] Object Management Group (2016): "Requirements Interchange Format (ReqIF)", Version 1.2.
OMG Document Number formal/2016-07-01.
[i.3] ISO/IEC Directives Part 2 (Edition 9) (2021): "ISO/IEC Directives, Part 2: 'Principles and rules for
the structure and drafting of ISO and IEC documents", Edition 9, 2021.
[i.4] Gotel OC, Finkelstein CW: "An analysis of the requirements traceability problem". In Proceedings
of IEEE international conference on requirements engineering 1994 Apr 18 (pp. 94-101). IEEE.
ETSI
8 ETSI TS 104 286 V1.1.1 (2026-08)
[i.5] Zhong R, Xu Y, Zhang C, Yu J.: "Leveraging large language model to generate a novel
metaheuristic algorithm with CRISPE framework". Cluster Computing. 2024 Dec;
27(10):13835-69.
[i.6] Almeida J. Prompt engineering: "A comparative study of prompting techniques in AI language
models". In 2025 IEEE Integrated STEM Education Conference (ISEC) 2025 Mar 15 (pp. 1-4).
IEEE.
[i.7] Andress J.: "The basics of information security: understanding the fundamentals of InfoSec in
theory and practice". Waltham, MA: Syngress. 2014.
[i.8] NIST SP 800-160 (Volume 1): "Engineering Trustworthy Secure Systems".
[i.9] NIST SP 800-53 (Revision 5) (2020-09): "Security and Privacy Controls for Information Systems
and Organizations". National Institute of Standards and Technology.
[i.10] Regulation (EU) 2024/1689 of the European Parliament and of the Council of 13 June 2024 laying
down harmonised rules on artificial intelligence and amending Regulations (EC) No 300/2008,
(EU) No 167/2013, (EU) No 168/2013, (EU) 2018/858, (EU) 2018/1139 and (EU) 2019/2144 and
Directives 2014/90/EU, (EU) 2016/797 and (EU) 2020/1828 (Artificial Intelligence Act),
OJ L, 12.7.2024.
[i.11] ETSI EN 303 645 (V3.1.3) (2024-09): "CYBER; Cyber Security for Consumer Internet of Things:
Baseline Requirements".
3 Definition of terms, symbols and abbreviations
3.1 Terms
For the purposes of the present document, the following terms apply:
digital transformation of standards: process of converting unstructured textual content of a standard into structured,
traceable, and machine-processable requirement records
hallucination: generation of content by a Large Language Model that is nonsensical, inaccurate, or unfaithful to
provided source material, yet presented in a confident and convincing manner
human-in-the-loop: AI development approach that integrates human interaction and expertise into the machine
learning cycle
requirement record: structured representation of an extracted requirement including its text, source reference, and
assigned security category
requirement statement: textual expression corresponding to a single identifiable security requirement extracted from a
standard
security category: predefined classification label representing a high-level security property to which a requirement
may belong
security requirement: normative statement within a standard that expresses an obligation, prohibition, or permission
related to security properties
traceability: ability to link an extracted requirement record to its original source location within a standard document
textual fragment: logically coherent portion of a standard document (e.g. clause, subclause, paragraph) used as input
for requirement identification and extraction
3.2 Symbols
Void.
ETSI
9 ETSI TS 104 286 V1.1.1 (2026-08)
3.3 Abbreviations
For the purposes of the present document, the following abbreviations apply:
AI Artificial Intelligence
BART Bidirectional and Auto-Regressive Transformer
BERT Bidirectional Encoder Representations from Transformers
CIA Confidentiality, Integrity, and Availability triad
CRA Cyber Resilience Act
DP Document Parsing
EU European Union
FP Feedback Provision
GPT Generative Pre-trained Transformer
IoT Internet of Things
JSON JavaScript Object Notation
LLM Large Language Model
MTS Methods for Testing and Specification
NLP Natural Language Processing
RC Requirements Classification
RCG Requirement Checklist Generation
RIE Requirements Identification and Extraction
XMI eXtensible Markup Language Metadata Interchange
XML eXtensible Markup Language
4 Background
Requirements are best represented not as isolated text fragments but as uniquely identifiable information objects. Each
object should include the normative requirement statement together with the metadata needed to support aspects such as
interpretation, provenance, classification, lifecycle management, rationale, verification method, and traceability.
Depending on the use case, this metadata can include a unique identifier, source artefact and location, type or
classification, version, and explicit relationships to other requirements and/or design elements, verification artefacts,
etc. [i.2].
The preservation of structural and contextual information is especially important when requirements are extracted from
standards or other normative documents. In such sources, the meaning of a requirement may depend on a lead-in
sentence, list structure, table context, or incorporated references. Consequently, requirement extraction has to preserve
not only the requirement statement itself, but also the surrounding structural context needed to interpret it correctly and
to reconstruct its normative intent [i.3].
Explicit representation of dependencies, references, and trace links is equally important. Traceability supports forward
and backward navigation across the lifecycle of a requirement, thereby enabling impact analysis, coverage analysis,
change management, and more reliable automation. A requirement representation that omits typed relationships may
still be readable to humans, but it is much less effective for rigorous analysis and digital processing [i.4].
Structured and machine-readable representations allow requirements, metadata, and relationships to be exchanged and
processed consistently across tools and organizational boundaries [i.2].
5 Methodology Overview
The input for the workflow described in the methodology is the complete document of a security standard (or
regulation) or fragments retrieved from the document, which is then analysed by AI models to semi-automatically
identify, extract, and classify security requirements that are present in the text of the analysed standard document. The
output is a list with the extracted security requirements in a human- and machine-readable format (see clause 6.6).
The workflow of the methodology is organized into three phases: a Parsing Phase, responsible for preparing the input
document for analysis; an Intelligence Phase, responsible for identifying, extracting, and classifying security
requirements; and a Reporting Phase, responsible for generating the resulting requirement checklist and incorporating
user feedback. The high-level overview of the methodology is illustrated in Figure 1.
ETSI
10 ETSI TS 104 286 V1.1.1 (2026-08)
Figure 1: The high-level overview of the proposed methodology
for the digital transformation of security standards
Each of the three phases consists of internal steps that collaboratively support the transformation process. The steps in
the workflow are described below:
• Parsing Phase
- Document Parsing (DP): This step constitutes the entry point of the workflow. It receives a standard or
regulation document and prepares its content for subsequent processing. This step produces structured
textual fragments together with associated contextual information (i.e. clauses and sections from which
the fragments originate) and provides them to the Requirements Identification and Extraction step.
• Intelligence Phase
- Requirements Identification and Extraction (RIE): This step receives the textual fragments that are
extracted by the DP step and analyses them to identify and extract security requirements. This step
produces structured and traceable requirement records containing the extracted requirements and
associated information (i.e. requirement identifier, clause and sub-clause in the standard). The output of
the RIE step is provided to the Requirements Classification and Requirement Checklist Generation steps.
- Requirements Classification (RC): This step receives the requirements identified by the RIE step in
order to classify each requirement to a security category. The output of the RC step consists of a set of
classified requirement items that preserve the traceability information provided by the preceding steps
and can be further processed by the RCG step.
• Reporting Phase
- Requirement Checklist Generation (RCG): This step receives the requirement records produced by
the RIE and RC steps and consolidates the extracted requirements, assigned security categories, and
associated information into a structured representation. The step produces a security requirements
checklist that constitutes the primary output of the digital transformation methodology.
- Feedback Provision (FP): This step enables users to inspect the produced requirements checklist and
provide feedback by validating, correcting, rejecting, or supplementing the requirements. The feedback
produced by the FP step is provided to the RIE and RC steps to support subsequent refinement of the
underlying models and configurations.
6 Methodology
6.1 Workflow
The workflow integrates DP, RIE, RC, RCG and FP steps. RIE and RC constitute the core analytical mechanisms, while
DP, FP and RCG steps provide supporting functionality. The requirements for these steps are outlined below:
a) The DP step:
a) shall act as the entry point of the digital transformation of standards pipeline;
ETSI
11 ETSI TS 104 286 V1.1.1 (2026-08)
b) shall transform a standard document into structured textual fragments suitable for processing by the RIE
step;
c) shall communicate with the RIE step by providing segmented textual fragments (e.g. clauses or
paragraphs) along with the clause and page identifiers (see clause 6.3).
b) The RIE step:
a) shall consume the textual fragments produced by the DP step and process them to identify the security
requirements contained in them;
b) shall preserve the clause and page identifiers associated with each textual fragment so that every
extracted requirement remains traceable to its corresponding location within the source document (see
clause 6.4).
c) The RC step:
a) shall process each requirement statement and assign a security category according to the predefined
categories list as described in clause 6.5;
b) shall pass the structured requirement records to the RCG and FP steps (see clause 6.5).
d) The RCG step:
a) shall aggregate requirement identifier, extracted text, corresponding clause and page, and assigned
security category in one record per requirement and generate machine-readable representations (e.g.
JSON, XML);
b) shall not modify semantic content when producing its output (i.e. requirement checklists) (see
clause 6.6).
e) The FP step:
a) shall receive the requirements checklist produced by the RCG and shall make it available for review by a
human expert;
b) shall enable the human expert to accept, reject, modify, or supplement the generated requirements. This
process shall support the correction of incorrectly extracted requirements, refinement of requirement
wording, adjustment of assigned security categories, and addition of omitted requirements based on
expert's judgement;
c) shall make feedback results available to the RIE and RC steps to be used as new data samples (see
clause 6.7).
This bidirectional communication establishes an iterative improvement loop that enhances the accuracy and stability of
both RIE and RC over time. It also reduces the risk of providing to the user output affected by LLM hallucinations.
6.2 Adaptation Techniques
The RIE and RC are based on the adaptation of pre-trained LLMs to the corresponding tasks. The adaptation technique
shall be selected according to the characteristics of the target application, the availability of labelled data, the
computational resources, and the expected performance. The methodology supports two major complementary
adaptation techniques:
• Prompt engineering: An adaptation technique for enabling the utilization of pre-trained LLMs for designated
tasks through prompting, without updating the parameters of the model via further training. Two types of
prompt engineering are recommended:
i) zero-shot learning where the prompt contains only instructions for performing the desired task without
providing any examples; and
ii) few-shot learning where the prompt contains both instructions and indicative examples. More details are
provided in clause 6.4.2.
ETSI
12 ETSI TS 104 286 V1.1.1 (2026-08)
• Fine-tuning: An adaptation technique for enabling the utilization of pre-trained LLM for a designated task
through updating its parameters via further training on a dataset that is specific to the desired task.
Prompt engineering enables the use of a pre-trained LLM without additional training, although its performance can
depend on the capabilities of the selected model and the formulation of the prompts. Fine-tuning enables the model to
be adapted to the designated task using task-specific data but requires labelled data and additional training, which can
be costly or infeasible due to the need for high-performance computing infrastructure.
Fine-tuning shall be preferred over prompt engineering when:
• smaller models need to be utilized for reducing the cost (e.g. energy consumption) at inference time;
• prompt engineering demonstrated insufficient performance for the desired standards or family of standards,
and therefore, the LLMs need to be fine-tuned to further improve their accuracy;
• focus is given on a specific standard or family of standards, and therefore, fine-tuning may be considered a
more viable option;
• the organization has the computational power required for fine-tuning LLMs.
The application of LLMs to the processing of standards can be subject to limitations such as the context-window size of
the selected model, loss of contextual information resulting from document segmentation, sensitivity to model and
prompt configurations, and availability of hardware resources. The methodology covers both adaptation techniques, in
order to cover a broad range of scenarios, stakeholders, use cases, and potential limitations.
6.3 Document Parsing (DP)
DP constitutes the first step of the digital transformation methodology and prepares the source document for subsequent
requirements identification and extraction. The DP step shall accept as input a security standard or regulation (or part of
it) in one of the various formats in which standard documents are commonly distributed (e.g. PDF, DOCX, XMI, etc.).
Then, the DP step shall split the document into a set of sequential textual fragments. Textual fragments may correspond
to clauses, subclauses, paragraphs, or other logical units of the source document.
The DP step shall pre-process the produced fragments to bring them into a form suitable for being processed by the
LLM-based requirements identification, extraction, and classification models. In particular, the relationship between
textual content and document structures relevant to the interpretation of requirements, including clause and subclause
boundaries, lead-in statements, lists, and associated list items, shall be preserved where present in the source document.
Where a textual fragment depends on preceding or subsequent content for its interpretation, sufficient context shall be
retained or associated with the fragment to support requirements identification and extraction.
The DP step shall associate each textual fragment with metadata sufficient to identify its location within the source
document. Such metadata shall include the document identifier, clause or subclause identifier, and page. The structured
textual fragments shall be provided as input to the RIE step.
6.4 Requirement Identification and Extraction (RIE)
6.4.1 Fine-tuning
6.4.1.1 Overview
Requirements identification and extraction can be achieved via fine-tuning of existing pre-trained LLMs. Fine-tuning is
based on the precondition that a reliable dataset containing textual fragments from standard documents, along with the
extracted requirements is available for use.
The high-level overview of the fine-tuning-based methodology is illustrated in Figure 2.
ETSI
13 ETSI TS 104 286 V1.1.1 (2026-08)
Figure 2: The high-level overview of the fine-tuning process for adapting
Large Language Models (LLMs) for security requirements identification
and extraction from standards text fragments
The fine-tuning process consists of five different steps:
1) Dataset Construction,
2) Data Preparation,
3) Model Selection,
4) Model Evaluation, and
5) Model Execution.
The first four steps are responsible for building the Requirements Identification and Extraction model, whereas the last
step corresponds to the usage of the produced model in practice. The main elements of this process are described in the
following clauses.
6.4.1.2 Dataset Construction and Data Preparation
The first step of the methodology is the construction of a domain-specific dataset, which is required for fine-tuning the
selected LLMs. The dataset shall consist of pairs of textual fragments (i.e. context) extracted from security standards or
regulations and the corresponding security requirement statements (i.e. sentences) contained within those fragments.
The textual fragments shall be paragraphs, clauses, or other logically coherent sections of the source standard.
The associated requirement statements shall correspond to the explicit security requirements expressed within the
corresponding context. The dataset construction process shall ensure consistency in annotation criteria across all
samples, traceability between each requirement statement and its source clause, appropriate coverage of different
requirement types and security domains.
The next step of the methodology is data preparation. Data preparation shall prepare the data for the model selection
process. The dataset shall be divided into training, validation, and testing sets. The training and validation sets shall be
used for fine-tuning the LLMs and selecting the best model, whereas the testing set shall be utilized for evaluating the
model's performance on unseen data samples. The data samples shall be tokenized to generate token (i.e. word)
sequences, which constitute the required format for the input of the LLMs.
ETSI
14 ETSI TS 104 286 V1.1.1 (2026-08)
6.4.1.3 Model Selection
The model selection step focuses on adapting pre-trained LLMs to the downstream task of security requirements
identification and extraction, using the domain-specific dataset described in clause 6.4.1.2. Model selection shall
determine an appropriate model configuration that can accurately identify and extract security requirement statements
from textual fragments of security standards. For this purpose, a variety of proprietary and/or open-source pre-trained
LLMs can be used, including Transformer-based architectures, such as BERT™, BART™, T5™, RoBERTa™, and
DistilBERT™. Pre-trained LLMs shall be fine-tuned using supervised learning techniques.
The expected output shall correspond to the associated security requirement statements. Therefore, it is considered a
Sequence-to-Sequence task, but different task formulations can be adopted for this adaptation process.
The pre-trained LLMs shall be fine-tuned on one of two objectives for the specific task of requirements identification
and extraction:
• Question-Answering objective;
• Summarization objective.
In Question-Answering, the requirement extraction task shall be structured as a question-answering problem. The model
shall:
• receive a textual fragment from a security standard (context);
• receive a predefined query prompt;
• be trained to generate the corresponding requirement statement or statements as the answer.
In Summarization, the requirement extraction task shall be structured as a targeted summarization problem. The model
shall:
• receive a textual fragment from a security standard;
• be configured to minimize abstraction;
• preserve the normative meaning of the original text, ensuring that extracted requirements remain semantically
aligned with the source content;
• produce a condensed output containing exclusively the security requirement statements present in the
fragment.
The selection of the most appropriate formulation and model configuration shall be based on evaluation results obtained
using the validation dataset. Performance shall be assessed using quantitative metrics such as ROUGE-L, precision,
recall, and F -score, as well as qualitative analysis of semantic fidelity and structural correctness. The selected model
shall demonstrate at least consistent identification of requirement statements across different clause structures and
reservation of the original normative intent.
6.4.1.4 Model Evaluation
During model's evaluation, the model configuration identified as best-performing during the validation process (i.e.
model selection step) shall be applied to unseen labelled data samples (i.e. the testing set). The evaluation results shall
be used to quantify the model's performa
...



