Between data and knowledge: the role of metadata in structuring information

Between data and knowledge: the role of metadata in structuring information

  • Duarte Dionísio Duarte Dionísio
  • 20 de may de 2025
  • 24 minutes

Metadata plays an essential role in the structuring, description and management of digital information, enabling data and content to be more easily found, interpreted, reused and preserved. Acting as ‘data about data’, metadata provides context and meaning to digital resources, ranging from documents and images to artificial intelligence models. Its main purpose is to ensure interoperability, the efficient organisation of information and its accessibility over time; it is therefore a central element in fields such as information science, library and information science, information technology and knowledge management.

Metadata is data about other data, functioning as structured descriptions that indicate the content, origin, context and other essential properties of the information. The Greek prefix "meta", meaning "beyond", illustrates this function well: metadata adds an extra layer of meaning that enables data to be organised, interpreted and reused effectively.

If we want a concrete example, let’s think of a book. The full text is the data itself. The title, author, publication date, subject or table of contents, on the other hand, are metadata. These elements are not part of the content itself, but describe it externally and in summary form, facilitating its cataloguing, searching and understanding. Without needing to read the entire content, it is possible to know what the work is about, who wrote it and where it fits in.

It is important to emphasise that metadata is neither to be confused with data nor does it replace it. It is a distinct layer of information that complements the content and enhances its use. In this sense, it functions as an invisible yet indispensable framework that guides the reading, analysis and management of information. It enables, for example, the verification of a file’s provenance, the assurance of its authenticity or the preservation of its integrity over time.

In a world dominated by digital information, metadata has become a kind of map that guides navigation through this ocean of data. Without it, much of the information would remain scattered, opaque or unusable. From historical archives to digital libraries, databases, search engines and artificial intelligence systems, metadata is now an essential cognitive infrastructure for transforming raw data into accessible and meaningful knowledge.

Importance and benefits

In an era characterised by an explosion of digital information, metadata plays a fundamental role in the management, retrieval and efficient use of data. Its relevance stems not only from the descriptive function it fulfils, but above all from its ability to transform large volumes of raw data into structured and accessible knowledge.

It makes it possible to locate information quickly and accurately. Without descriptive and structural metadata, searching for a specific resource amongst thousands of records would be time-consuming and, often, fruitless. Assigning attributes such as author, title, keywords or date of creation enables search engines to filter, index and classify content based on defined criteria. In a digital library, for example, simply linking a book to its author and subject makes it possible to narrow down search results to those items that are truly relevant.

Furthermore, metadata significantly improves the internal organisation of databases. It not only describes the data, but also defines how the data relates to one another. In relational systems, for example, it is the metadata that determines the structure of tables, primary and foreign keys, and the links between datasets. Without this structural layer, operations such as complex queries or joins would lose efficiency, consistency and reliability. Metadata thus functions as the grammar that governs the language of data; without it, communication between elements becomes chaotic or impossible.

Another benefit relates to interoperability and integration between different systems. In an increasingly interconnected digital ecosystem, it is vital that different applications are able to understand and reuse data produced on other platforms. Metadata, when organised according to common standards and schemas (such as Dublin Core or DCAT-AP), ensures this semantic and technical compatibility. For example, public datasets from various countries can be integrated into a single portal, provided they use shared metadata standards. Similarly, national libraries are able to exchange records with one another thanks to the adoption of standardised descriptors.

In this sense, metadata is the invisible infrastructure that underpins communication between systems, facilitating everything from simple tasks, such as searching for documents, to more complex applications, such as data integration in big data environments, the automation of administrative processes, or even the training of artificial intelligence models. It is, therefore, one of the central pillars of any open, sustainable and effective data strategy.

In Portugal, the National Register of Digital Objects (RNOD) centralises bibliographic metadata from various institutions, aligning it with Europeana’s international standards and publishing a large volume of records since 2013. This effort ensures high quality and seamless interoperability with the European Union’s digital library. The Portuguese Open Access Scientific Repository (RCAAP), in turn, aggregates descriptions of articles, theses and reports deposited in national institutional repositories (and, via OASIS, in Brazilian repositories), creating a meta-repository that makes thousands of publications accessible solely through their metadata, without storing the full text. Finally, the European Data Portal adopts the DCAT-AP profile to standardise the descriptions of public datasets from various countries.

Metadata in different contexts

The application of metadata varies widely depending on the field of use, playing a critical role in the organisation, retrieval and interpretation of information across a range of scenarios.

Libraries and archives

Library science emphasises the efficient management of physical and digital collections. Cataloguing systems, such as MARC21 or Dublin Core, use descriptive metadata to organise books, journals, documents and cultural artefacts. Fields such as title, author, subject and publication date make it possible to quickly locate relevant resources, whether in physical catalogues or digital libraries. Its function goes beyond simple searches: it helps to structure and categorise collections, facilitating the discovery of related materials and coherent navigation. Without structured metadata, digital libraries would become disorganised repositories, hindering efficient access to information. The versatility of standards such as Dublin Core is evident in its ability to describe both a physical book and a web page using the same fields.

Web and SEO

In the context of the internet, metadata is embedded in HTML code in the form of meta tags (such as title, description or keywords) and semantic markup (such as JSON-LD). Search engines need this metadata to understand, index and rank the content of web pages. Defining accurate and relevant metadata significantly improves online visibility by promoting better ranking in search results (SEO). The use of standardised vocabularies such as Schema.org enriches web pages with markup that is understandable to search engines and social media platforms, which helps drive qualified traffic to the website. In a highly competitive online environment, metadata is essential for ensuring that content reaches the right target audience.

Corporate and scientific databases

In business and academic contexts, metadata provides a structured description of how databases are organised, specifying the fields, the relationships between tables and the rules that ensure the integrity of the information. Specifications such as primary and foreign keys enable information to be integrated and analysed consistently and efficiently. Furthermore, data lineage metadata documents the origin and journey of data throughout its lifecycle, which is crucial for ensuring confidence in results and enabling audits. In scientific repositories, for example, metadata makes it possible to search for articles, datasets or reports by author, keywords or persistent identifiers such as the DOI. In this context, where it is crucial to verify results, metadata records the path of the data and the operations performed, enabling studies to be reproduced accurately.

Multimedia and digital assets

In the field of photography and video, each digital file contains metadata such as EXIF, IPTC or XMP, which describe technical aspects (such as date, device and geolocation) and contextual aspects (such as authorship or usage rights). Organisations that manage large volumes of multimedia content rely on digital asset management (DAM) systems, where metadata is essential for organising, searching for and distributing files effectively. This metadata not only makes it easier to locate assets, but also ensures that usage rights are respected and that content is reused appropriately.

Document management

In business environments, document management requires metadata to ensure version control, the identification of changes and the definition of access permissions. Metadata enables the identification of the author, the creation date, retention periods and the access policy for each document. This control is crucial for ensuring compliance with regulations and audits, as well as for safeguarding the integrity and security of information.

E-commerce and content personalisation

In e-commerce, metadata tracks user behaviour — clicks, purchase history, preferences — and feeds into recommendation algorithms. This data makes it possible to personalise the customer experience and optimise logistics management, contributing to greater operational efficiency and increased end-user satisfaction. Well-structured metadata makes it possible, for example, to suggest complementary products, anticipate stock replenishment needs or assess seasonal consumption patterns.

Types of metadata

In practice, there are several main categories of metadata, each with a different function:

Descriptive metadata

This refers to attributes that identify and characterise the resource, facilitating its discovery and recognition. It provides essential information that helps users locate what they are looking for. Traditional examples include the title of a document, the author of a work, the subject, the date of creation and keywords that summarise the content, as well as the summary or abstract. These fields simplify the retrieval and classification of documents (for example, in a library catalogue, where they allow books to be catalogued by title, author or genre, making searches more efficient; or on a social network, where they guide search engines in the correct categorisation of resources available on the web). They form the basis of information discovery, acting as labels that enable content to be identified and classified, and serve as the gateway to accessing information.

Structural metadata

This describes how data elements are organised internally and defines the relationships between their components. It indicates how the different elements of an object are interconnected. For example, it specifies how web pages or sections of a document relate to one another in a hierarchy or sequence, such as the order of chapters in an e-book. In relational database systems, this metadata specifies the links between tables. By mapping the internal relationships of complex datasets, structural metadata facilitates an understanding of their architecture and organisation, acting as the skeleton of the information and showing how the various parts fit together to form a cohesive whole.

Administrative metadata

This involves management information and operational context, and is responsible for managing access to a resource, defining permissions and controlling the different versions of a document or dataset. This includes data such as who created the object, who can access or modify it, modification dates, usage rights (often represented by a URL linking to a licence document), file formats or retention policies. These fields support access control, preservation and legal compliance of the data, and are particularly important for ensuring regulatory compliance. Administrative metadata can be subdivided into categories such as technical metadata, preservation metadata (which covers the archiving and long-term management of assets, and may include a checksum to ensure their integrity) and rights metadata (which indicates intellectual property and usage rights, e.g. «© 2025 Duarte Dionísio. All rights reserved.»). It is crucial for data management and security, ensuring that resources are managed appropriately and in accordance with policies and regulations.

Technical metadata

This describes the technical properties of the resource, capturing the "physical" characteristics of the data, such as file type and format, size, resolution, compression parameters and encodings, amongst others. In photographic records, for example, it includes details of the camera used, settings and the location where the photograph was taken. In data warehousing and business intelligence (DW/BI) systems, it defines objects and processes from a technical perspective, including data structures such as tables and fields. This metadata supports the proper rendering of data and facilitates processes; it is essential for ensuring the correct processing and interpretation of data, as well as ensuring compatibility between systems.

In summary, each category of metadata offers a complementary view of the resources, making it possible to organise, access and preserve information effectively.

Standards and specifications

The adoption of standardised schemas is crucial to ensuring that metadata remains consistent, interoperable and understandable across different systems and organisations. Below is a compilation of the main standards, indicating the body responsible for their maintenance and, where applicable, the associated ISO standard or equivalent.

Dublin Core

Dublin Core, maintained by the Dublin Core Metadata Initiative (DCMI), defines 15 basic elements for describing any digital or physical resource:

  • Title – Title of the resource.
  • Creator – Author or main person responsible for creating the content.
  • Subject – Theme, subject or keywords associated with the content.
  • Description – Description of the content (summary, synopsis, etc.).
  • Publisher – Organisation responsible for making the resource available.
  • Contributor – Other individuals responsible for contributing to the creation of the content.
  • Date – Date associated with the resource (creation, publication, etc.).
  • Type – Type or nature of the resource (text, image, sound, etc.).
  • Format – Physical or digital format of the resource (MIME type, file extension, etc.).
  • Identifier – Unique identifier of the resource (such as a URL, DOI, ISBN, etc.).
  • Source – Original resource from which this one derives, if applicable.
  • Language – Language of the content (e.g. pt, en, fr).
  • Relation – Relationship with other resources (for example, a previous or associated version).
  • Coverage – Spatial or temporal scope of the content (for example, geographical location or historical period).
  • Rights – Information on the resource’s usage rights and intellectual property rights.

Its simplicity and flexibility make it suitable for digital libraries, web pages and academic repositories alike. In 2003, Dublin Core was formalised as ISO 15836-1, establishing it as an international standard.

MODS

Developed by the US Library of Congress, MODS (Metadata Object Description Schema) is an XML schema that adds granularity to Dublin Core, incorporating detailed fields such as notes, editions and series. Widely used in conversions between MARC21 and web formats, it is valued by libraries requiring more detailed descriptions and is officially maintained by the Library of Congress.

EAD

EAD (Encoded Archival Description), created jointly by the Society of American Archivists and the Library of Congress, is an XML standard dedicated to the description of archival fonds and collections. It allows complex hierarchies (fonds, series, record) to be represented, whilst preserving the historical context of the records. It originates from the North American archival community.

ANSI/NISO Z39.87

This ANSI/NISO Z39.87 standard (Technical Metadata for Digital Still Images), developed by the National Information Standards Organisation in partnership with the Library of Congress, defines technical metadata for still images — resolution, colour profile, capture equipment and compression parameters. It is essential for ensuring the uniform interpretation of the properties of digital images.

MIX

MIX (Metadata for Images in XML) complements Z39.87 by translating this technical image metadata into XML, facilitating automatic validation and integration into digital preservation workflows. It is regarded as the leading practical standard for image management in libraries and repositories, and is maintained by NISO.

METS

METS (Metadata Encoding and Transmission Standard), developed by the Library of Congress and the Digital Library Federation, acts as an XML container that aggregates descriptive, administrative and structural metadata for digital objects. It can incorporate schemas such as Dublin Core, MODS or MIX and describes in detail the physical structure (pages, components) and logic (order, versions) of complex resources. METS is managed by the METS Editorial Board under the guidance of the W3C.

PREMIS

PREMIS (Preservation Metadata: Implementation Strategies) is a data dictionary focused on long-term digital preservation. Maintained by the PREMIS Editorial Committee, it defines entities (objects, events, agents) and records preservation operations, such as migrations, file integrity (checksums) and changes to rights.

Schema.org

A W3C-supported recommendation, created jointly by Google, Microsoft, Yahoo and Yandex, Schema.org is a collaborative vocabulary for the semantic marking-up of web pages. Using microdata, RDFa or JSON-LD, it enables entities (products, events, people, organisations) to be annotated in a way that is understandable to search engines, thereby optimising visibility and SEO.

FOAF

FOAF (Friend-of-a-Friend) is a pioneering RDF/OWL ontology in the Semantic Web, promoted by the W3C community. It describes people, activities and social relationships in a decentralised manner. It enables the creation of interoperable public profiles without a central database.

ONIX for Books

ONIX for Books, maintained by EDItEUR, establishes an XML format for the exchange of bibliographic and commercial metadata about books — including title, author, ISBN, price, format and rights. ONIX is the preferred format for sharing data about books and publications, as a result of widespread adherence to EDItEUR’s recommendations.

IPTC

The IPTC Photo Metadata Standard defines a structured set of fields for describing digital images, organised into the IPTC Core and IPTC Extension schemas. These range from descriptive and rights-related information to technical and workflow metadata, and are widely adopted in communications, archives and DAM systems. Since 2004, the standard has been compatible with the XMP format, facilitating integration with tools such as Adobe Photoshop and promoting interoperability with other standards such as EXIF.

Creation and management tools

There are several categories of tools that support the creation, editing and maintenance of metadata. Cataloguing systems, such as Koha, Aleph or Biblivre, integrated into library management systems (ILS), enable the generation of bibliographic records with standardised metadata, such as MARC21, including authority control and export via OAI-PMH or APIs. Content management systems (CMS), such as WordPress, Drupal or ConcreteCMS, and digital asset management (DAM) platforms, such as ResourceSpace or Canto, incorporate descriptive and technical metadata fields, supporting vocabularies such as Schema.org, thereby optimising the management of images, videos and documents with tags indicating content, format or usage rights.

Furthermore, validation and editing tools, such as JSON-LD, RDFa and XML/XSD validators, or editors such as Oxygen XML Editor, Protégé and OpenRefine, ensure that metadata complies with schemas such as Dublin Core or Schema.org, enabling the detection of structural errors. Data repositories and catalogues, such as DSpace, EPrints, Fedora Commons or CKAN, are essential for organising and making digital collections or public datasets available, often using standards such as DCAT-AP. When used correctly, these tools automate tasks, improve the quality of metadata and make it easier to update, thereby reducing the effort required for indexing and information retrieval processes.

Ethical and legal considerations

The processing of metadata poses a significant risk to individuals’ privacy, as it may reveal sensitive information even when the main content remains protected. Communications metadata, such as the sender, recipient, date and location, is highly confidential and must be anonymised or deleted whenever it is not strictly necessary. The application of the General Data Protection Regulation (GDPR) in the European Union and Law No. 58/2019 in Portugal also extends to personal metadata, imposing obligations on organisations regarding encryption and strict access control to ensure legal compliance and protect the privacy of data subjects.

On the other hand, well-structured and enriched metadata enhances digital accessibility. Recent standards, such as the IPTC standard for image metadata, already include fields such as "Alt Text" and "Extended Description" embedded within files, enabling screen readers and other assistive technologies to present the content in a way that is understandable to people with visual impairments. Directive (EU) 2016/2102 and the corresponding national Decree-Law 83/2018 require public administration websites to provide accessible information; this also involves including text descriptions for images, videos and other multimedia resources, ensuring that all users have equal access to the content.

Copyright also requires that metadata clearly records the licences and terms of use for content. It is common practice to use fields such as "rights" and "licence" in the Dublin Core schema to indicate whether a resource is subject to copyright, Creative Commons or another permissions model. Keeping this data permanently up to date prevents misuse and potential legal disputes by providing users and information systems with a clear and reliable reference regarding the conditions for reproducing and sharing protected content.

Alongside privacy and security requirements, the Portuguese Parliament has been considering the creation of a legal framework for the retention of electronic communications metadata. In 2022, Draft Law No. 11/XV/1st required telecommunications operators to retain metadata, such as telephone numbers, IP addresses and device identifiers, for up to six months, without the need for prior authorisation from users. However, the Constitutional Court deemed this blanket retention to be disproportionate and unconstitutional, and rejected it on the grounds that it violated the rights to privacy and to informational self-determination. In response to this ruling, a new approach was adopted and, in January 2024, legislation was passed which became Law No. 18/2024 of 5 February, regulating access to metadata for the purposes of criminal investigation. This new framework makes the retention of traffic and location metadata subject to a reasoned judicial authorisation, to be issued within 72 hours, and solely for the investigation, detection and prosecution of serious crimes. This law ensures not only respect for our privacy, but also the importance of keeping metadata – which may be decisive in criminal proceedings – available, subject to strict and scrutinised control, thereby guaranteeing the integrity of the judicial system.

Challenges in metadata management

Despite the obvious advantages associated with the use of metadata, managing it presents challenges that require strategic planning and well-informed technical decisions. One of the most pressing problems is inconsistency: when different users describe the same content in divergent ways, using different vocabularies or omitting relevant fields, the result becomes chaotic, difficult to search and unreliable. Automating validation and updating processes mitigates these errors, but this always depends on the existence of clear standards that are properly communicated within organisations. At the same time, scalability necessitates infrastructure capable of supporting the creation, storage and processing of metadata on a large scale; without scalable solutions, performance deteriorates. Management emerges as a third critical pillar: without organisational policies defining who creates, reviews or approves metadata, redundancies, inconsistencies and a loss of control over records proliferate. Added to this is constant technological evolution, in which formats and schemas quickly become obsolete – risking rendering years of cataloguing work useless – and which requires strategic planning.

At the same time, emerging technologies promise to redefine the role of metadata. Artificial intelligence and machine learning require well-structured data for their models: the more complete and consistent the metadata is, the more efficient and reliable the automatic generation of tags, summaries or entity extraction will be. Blockchain is emerging as a mechanism designed to guarantee the immutability and authenticity of records, generating digital certificates that document the provenance and authorship of content. In the Semantic Web and Linked Data, RDF/OWL ontologies (such as FOAF) enable concepts to be linked intelligently, paving the way for repositories that interconnect automatically. The Internet of Things, which generates massive volumes of records, only becomes meaningful when accompanied by standardised metadata, such as timestamps, GPS coordinates and sensor types, and enables these devices to publish their own semantic context.

Metadata has become just as indispensable as the data itself, serving as the cornerstones of research, organisation and interoperability in large-scale digital environments. However, crucial questions remain: how can we ensure the quality and consistency of automatically generated metadata? How can we safeguard privacy and ethical standards when collecting and sharing data across distributed networks or IoT devices? Is it feasible to reconcile interoperability with the multitude of standards and institutional policies? And how can we design flexible management models that keep pace with rapid technological change, without compromising reliability or security?

 

References and inspiration