Chaos Computer Club - archive feed

Chaos Computer Club - archive feed

By CCC media teamTechnology
Download on the App Store

Chaos Computer Club - archive feed episodes

  • Deploy software with systemd-sysext (froscon2022)
    Various container runtimes exist on Linux to run software without installing packages on the host system. The nature of containers implies separation of the host system which sometimes is a gap that needs to be bridged again. For systemd services the "portable service" format allows to run a service with its own dependencies bundled in a filesystem image. However, like a container it still does not make any CLI tools directly available to the host system. Therefore, a common solution is to copy a set of static binaries to the system to use the same deployment mechanism for the service and the CLI tools. The new systemd-sysext format allows to extend the host system through an overlay that integrates the bundled software similar as traditional packages do. The binaries and config files can be updated and managed through a single sysext image file. A version matching logic allows to ensure that a particular host system version is used for depending on certain features or for dynamical linking. We demonstrate how systemd-sysext helps to extend an immutable host system such as Flatcar Container Linux, both for third party user software as well as an internal building block for more modularity.
    Linux package managers provide a mature way to install, update, and remove additional software on a system. If using packages is not possible or wanted, containers get used, e.g., through Docker. Containers come with upsides, such as reduced dependency requirements and increased isolation, but also downsides because they don't integrate with the system as well as packages would do. There are workarounds like creating systemd services that start the container and expose expected APIs to the system and using a wrapper script to make the container behave like the CLI tool does if installed through a package. Another approach is to use statically linked binaries for the service and the CLI tools or at least for CLI tools to complement a container. The systemd project introduced support for “portable services” which addresses the integration of a service with the system. It provides a way to set up systemd services from a container-like filesystem image that contains the systemd service definition and the required binaries and dependencies. The recent systemd-sysext format aims to address the extension of the system with additional CLI tools. It works by managing overlay mounts of the sysext images on top of the “/usr” and “/opt” system folders. This has the benefit of bundling a set of binaries inside a single image file that can be added, updated, and removed atomically. There is also a version matching logic that enables safe usage of dynamic linking and depending on certain features of the OS by ensuring that a particular host system version is used. While the primary use case is for deploying additional tools, it can also be used to provide systemd services and their binaries or temporarily overwrite host system files at runtime. Using systemd-sysext for extending the system with additional systemd services requires a small workaround but allows to bundle a service and its CLI tools into a single sysext image. In result, it integrates well with the system and behaves similar to software installed through a regular package.
    There are many use cases possible for sysext images, and we will demonstrate some of them for Flatcar Container Linux, which has no package manager but at the same time needs to be open for user customization, selection of container runtimes, and optional cloud vendor tools.
    about this event: https://programm.froscon.org/2022/events/2775.html
    44 min
  • batou (froscon2022)
    batou ist ein Python-basiertes Deployment-Werkzeug, das sich für einfache und komplexe Anwendungen eignet. Im Vortrag werden die wesentlichen Grundideen hinter batou erläutert indem wir ein Deployment für die Konfiguration von Cumulus-Switches bauen.
    Eine besondere Eigenschaft von batou stellt dabei das fraktale Modell dar: es erlaubt den leichten Wechsel zwischen deklarativer Modellierung und imperativer Implementation, sodass man sich auch in komplexen Situationen auf Eigenschaften wie Idempotenz, Konvergenz, Konsistenz und Vorhersagbarkeit verlassen kann.
    about this event: https://programm.froscon.org/2022/events/2817.html
    1 hr 1 min
  • Opencanary, eine Alarmanlage gegen Einbrecher im Netzwerk (froscon2022)
    Schaffen Sie mit OpenCanary eine Netzwerkalarmanlage, mit dem Sie Hacker abfangen können, bevor diese Ihre Systeme vollständig kompromittieren.
    Unternehmen brauchen in der Regel 6 Monate, um herauszufinden, dass sie Opfer einer Cyberattacke waren. Je länger es dauert, desto kostspieliger wird der Vorfall. Wenn man sich mit den digitalen Angriffen und Reaktionsempfehlungen beschäftigt, ist man als Organisation mit einer kleinen IT Abteilung schnell überfordert.
    Immer wieder liest man dann
    - Logdateien analysieren
    - Netzwerkverkehr überwachen
    - Intrusiondetection Software installieren
    Alles ist richtig, aber die Komplexität und die Kosten für solche Maßnahmen sind kaum zu stemmen. Hier kann die Opensource Software Opencanary helfen.
    Einfach gesprochen schafft OpenCanary einen Netzwerk-Honeypot, mit dem Sie Hacker abfangen können, bevor sie Ihre Systeme vollständig kompromittieren. Technisch bedeutet das: OpenCanary ist ein Daemon der Alarm schlägt, wenn ein Dienst (miss)gebraucht wird. Man kann dann festlegen, dass man zum Beispiel einen FTP Server, einen Fileserver oder einen Datenbankserver simulieren möchte und wohin eine Alarmierung gesendet werden soll.
    In dieser Session betrachten wir, was passieren sollte, wenn sich ein Angreifer im Netzwerk umschaut und versucht weitere Services anzugreifen.
    about this event: https://programm.froscon.org/2022/events/2732.html
    1 hr 3 min
  • Returning the favor - Leveraging quality insights of OpenStreetMap-based land-use/land-cover multi-label modeling to the community (sotm2022)
    The fitness of OSM for multi-label classification is proven. A workflow to enhance OSM-based multi-labels using machine learning is established. The results are provided to the OSM community via the HOT Tasking Manager.
    # Introduction
    Land-use and land-cover (LULC) information in OSM is a challenging topic. On the one hand, this information provides the background for all other data rendered on the central map and is used by applications like https://osmlanduse.org. On the other hand, this information has a difficult position within the OSM ecosystem. LULC information can be quite cumbersome or even difficult to map e.g. due to natural ambiguity. The growing tagging scheme provides a collection of sometimes ambiguous or overlapping tag definitions that are not fully compatible with any official LULC legend definition [1]. Furthermore, the data is highly shaped by national preferences and imports.
    This diversity of the LULC data in OSM is a fundamental principle of OSM that enabled the success of the project. Yet, this can create considerable usage barriers or at least caveats for data users unfamiliar with the projects' ecosystem. The remote sensing community for instance has started to use OSM LULC information as labels in their classification models. Frequently, OSM LULC data has thereby been taken at face value without critical reflection. And, while the quality and fitness for purpose of OSM data has been proven in many cases (e.g. [2,3]) these analyses have also unveiled quality variations e.g. between rural and urban regions. The quality of OSM therefore can be assumed to be generally high, but remains unknown for a specific use-case.
    The proposed work first assesses the impact of these challenges on a use-case of multi-label remote sensing (RS) image classification and then provides a machine learning (ML) based workflow to overcome and finally mitigate them. Multi-labels are a type of image classification where a satellite image is labeled with multiple containing LULC classes. In the presented study these labels are extracted from OSM and used to train the ML algorithm.
    # Methods and Results
    The fitness for purpose of OSM for multi-label RS image classification was tested on a Sentinel 2 scene with a resolution of 10m and four bands in south west Germany recorded in June 2021. The area was chosen for its estimated high completeness and low amount of imported data. OSM data was grouped by its tags into the four LULC classes 'forests', 'agricultural areas', 'build-up area' and 'water bodies'. 18 tags that could unequivocally be mapped to these classes were used and small elements below the image resolution or the classes minimal mapping unit were filtered. The chosen scene was then tiled into a 1.22 x 1.22 km grid of 8100 image patches. Zero to four labels were assigned to each patch, based on the OSM LULC elements therein. Evaluation was performed manually on 910 random patches, of which 80% had a correct OSM-based multi-label, thereby proving the assumed high completeness and quality in the region.
    The proposed workflow provides a method to enhance this OSM-based RS image multi-label classification and extend it to areas of lower OSM quality and completeness using ML (specifically deep learning (DL)). The main obstacle for ML and especially DL is the required amount of labeled training data. Volunteered geographic information (VGI) like OSM offers a potential solution to this challenge by providing an overabundance of LULC information that is suitable for this purpose if data quality is sufficiently high. The workflow uses the multi-label information extracted from OSM for training and then detects discrepancies between its predictions and OSM.
    Using this information and pinpointing the exact location of error within the patches provides valuable OSM data quality information. Apart from facilitating a fast quality estimation for large areas, the workflow can make its findings automatically available to the OSM community in a feedback loop using the HOT Tasking Manager framework. Thereby the valuable service by the OSM community of providing large amounts of free and generally high quality training data is recognised in the form of quality feedback including mapping hints to the OSM community.
    The five workflow stages are: 1) RS data collection and preprocessing, 2) OSM data collection and preprocessing, 3) LULC multi-label modeling, 4) OSM data issue flagging and 5) the community feedback loop. While each step is an atomic use case and application, the combination of all four steps creates a tool that is useful for the RS and the OSM community likewise. The tool is openly available at https://gitlab.gistools.geog.uni-heidelberg.de/giscience/ideal-vgi/osm-multitag under the GNU Affero General Public License v3 including example datasets. Manual input was kept as low as possible while enabling the 'human in the loop' to take full control over all input and output.
    The workflow extracts multi-label training data in stages 1) and 2) as described. Stage 3) then trains a DL model to predict multi-labels using solely RS imagery. For demonstration, the model was trained on the described Sentinel 2 scene in Germany. The models' performance was validated on the manually labeled 910 patches where it outperformed OSM in terms of multi-label accuracy by 7%. When additional errors were manually introduced to the training labels to simulate areas of lower OSM quality or completeness, the model maintained an overall prediction accuracy above the noisy training labels. Alternatively, in cases where OSM LULC multi-label accuracy is suspected to be low, pretrained models from comparable regions with higher OSM data quality can be applied, making the workflow widely applicable.
    Any patches where the models' multi-label prediction contradicts the OSM-based multi-label are then detected in stage 4). Multi-labels can be incorrect if either a label is missing (omission), meaning data is missing in OSM, or if a label is wrongly assigned, meaning OSM data is falsely mapped within the tile. The data error type and location within the patch is then extracted using explainable AI [4].
    The final stage 5) uses these localised potential OSM data errors to create HOT Tasking Manager projects via the public API. These projects provide additional correction hints. Yet, no automatic editing takes place. The community is kept in full control of all mapping actions as a 'human in the loop'.
    # Discussion
    The high quality but diverse nature of OSM has been proven for the use-case of multi-label RS image classification. The proposed tool provides an automated OSM multi-label extraction, modeling and verification procedure including a return of results to the OSM community. A major challenge of the approach is the tiled view on the data. If OSM assigns correct multi-labels to a patch, more fine grained data issues will not be detected. Yet, this approach allows large scale data assessments, before the topic of more detailed data improvement is tackled. It also allows to run repeated OSM LULC quality and completeness estimations for large areas over time.
    Another major benefit is the usage of local OSM data for modeling, thus making regionalised models the standard procedure. This is required for OSM LULC information as regional data structures and communities exist, that need to be preserved. The model can lead to regional homogenisation and data cohesion within these regional communities.
    about this event: https://2022.stateofthemap.org/sessions/EKEZ7R/
    6 min
  • Corporate editing and its impact on network navigability within OpenStreetMap (sotm2022)
    Using intrinsic quality indicators we explore how network quality, in terms of its suitability for navigation, varies across areas with relatively high and low corporate editing in OpenSteetMap. Our work shows areas with relatively high rates of corporate editing exhibit not only an overall increase in data quality, but also increased rates at which quality improves.
    OSM (OSM) contributors have traditionally lacked explicit monetary incentives for contribution [1]. Since 2016, a handful of large corporations (including Apple, Facebook, Microsoft, and Uber) have increasingly contributed data to OSM. Corporate editors (CEs) represent a distinct community as their editors are compensated and thus their contributions cannot be labeled as ‘volunteered’. Additionally, corporations employ large editing teams and new state-of-the-art editing techniques aided by artificial intelligence, making them capable of editing large swaths of information in relatively short time [2]. Corporate teams are often led by long-time OSM community members themselves, emphasizing the multifaceted nature of a rapidly growing open mapping platform [3]. While there has been some contention about the quality of edits done by CEs, corporations argue their contributions improve existing data [7]. Our study provides a preliminary quantitative evaluation of data quality impacts of corporate edits on OSM.
    We assess intrinsic data quality across five regions that have high levels of corporate contributions: Dallas-Ft. Worth, Egypt, Jamaica, Thailand, and Singapore. The quality of these regions is compared to that of Denmark, a region which has witnessed relatively less corporate interest, yet possesses a well-mapped OSM presence due to a well developed local mapping community [4]. These evaluations were performed using measures of intrinsic map quality. While the most straightforward evaluation methods involve comparing against extrinsic sources, such as either ground reference information or authoritative data sources; lack of data availability, licensing terms, and costs often render this comparison untenable [5,6,7]. A transferrable, data driven way of assessing quality remains using Intrinsic Quality Indicators (IQIs), a sub-field of OSM analysis which provides a variety of approaches for evaluating intrinsic OSM data quality. We chose to focus on IQIs that apply to networks, and to evaluate IQIs for land-based transportation networks within OSM. We analyzed networks for our specified locations for every other year between 2014 and 2022.
    OSM editing archives were processed using R to extract maps of the relative activity of corporate editors [8]. Our list of corporate editors was sourced from OSM’s publicly available list of corporate editors accounts. We extracted entire networks that represented the first day of each year of interest (2014, 2016, 2018, 2020, 2022) from OSM’s historical archives. For the purposes of this study, we extracted all networks where “OSM WAY = Highway”.
    We evaluated several IQIs for our areas of interest. We focused on completeness of network, both in terms growth over time and in terms of its navigability. We operationalized “completeness for navigability” as an intrinsic measurement by exploring the percentages of networks that possessed attributes necessary for GPS navigation – street names and speed limits. Navigability was assessed and compared across time points using Origin-Destination matrices. By creating a regular matrix across the area and calculating the ratio between a direct route between points, and a route navigated within our network, we calculated a ratio that can be compared across time to evaluate the changing efficiency of the navigable network. Additionally, when building routing networks, we discovered an additional IQI : the presence and qualities of topological islands within our network. That is, areas which are disconnected from the main network due to mapping errors or incompleteness.
    After mapping these metrics, we analyzed how they correlate with each other and how they change over time. Overall, IQI trends for the road network reveal consistent patterns across all measures and locations. There is a trend towards increasing data quality in terms of gradual increase of network length, completeness in terms of attributes (name, speed limit, and pedestrian access), the increasing efficiency of ODM routing ratios, and the increasing amount of places that have “navigable” attributes. Importantly, we found differences between our control location (Denmark) and our other areas of interest. The primary difference of note is not with regards to the quality of the data, but with respect to the rate at which data quality improves: Denmark’s rate of quality improvement is slower than other locations. The faster rate of quality improvement in the test areas highlights that the data creation and editing activity by corporate editors and other organized editors in these locations are helping narrow gaps in data quality.
    While this presentation highlights the trends of data quality increase, it does not tease apart the quality assessment of contributions by corporate teams versus other mapping groups. As a crowdsourcing platform, data in OSM is co-produced by repeated editing of data objects by different members of the community [9]. The appearance of CEs in OSM represents the arrival of another community of ‘produsers’ in the OSM ecosystem, and thus a new evolution in its overall trajectory [2,10,11]. Consequently, there is significant interaction between CEs and non-CEs in data co-production in OSM, further reinforcing the idea that OSM is a ‘community of communities’ [11]. Each location has their own patterns regarding editing communities, what they edit, and the sociopolitical and economic ground truth in the real world. Each of these factors impact the data, and may make comparing editing patterns difficult, especially given the diversity of motivations both with CE communities and within other OSM communities. Hence, we do not try to pry apart the differences in trends between individual countries. Instead, we focus on the overall trend between our test and control locations. With these caveats, we find that the quality of the network has increased in these areas across all tracked metrics at a faster rate than it has in areas with low rates of corporate edits, indicating that corporate editing may have a positive effect on the overall quality of the map.
    about this event: https://2022.stateofthemap.org/sessions/EZPVPB/
    8 min
  • Investigating the capability of UAV imagery in AI-assisted mapping of Refugee Camps in East Africa (sotm2022)
    This pilot project is connected to a larger initiative to open-source the assisted mapping platform for Humanitarian OpenStreetMap (HOTOSM) based on Very High Resolution (VHR) drone imagery. The study test and evaluate multiple U-Net based architectures on building segmentation of Refugee Camps in East Africa.
    Introduction
    Refugee camps and informal settlements provide accomodation to some of
    the most vulnerable population, the majority of which are located in Sub-
    Saharan East Africa (UNHCR, 2016). Many of these settlements often lack
    up-to-date maps of which we take for granted in developed settlements. Hav-
    ing up-to-date maps are important for assisting administration tasks such as
    population estimates and infrastructure development in data impoverished
    environments, and thereby encourages economic productivity (Herfort et al.,
    2021). The data inequality between the developed and developing countries
    are often resulted from a lack of commercial interest, especially with the
    recent trend of corporate OSM mappers (Anderson et al., 2019, Veselovsky
    et al., 2021). Such disparity can be reduced using assisted mapping tech-
    nology. To extract geospatial and imagery characteristics of dense urban
    enviornments, a combination of VHR satellite imagery and Machine Learn-
    ing (ML) are commonly used (Taubenböck et al., 2018). Classical ML based
    methods that exploit the textual (e.g. GLCM), spectral, and morphological
    characteristics of VHR imagery are based on the principles of Computer
    Vision (CV). Although many have shown promising results in satellite VHR
    (1m to 5m resolution) scenarios such as differentiating slum and non-slum
    (Kuffer et al., 2016 & Wurm et al., 2021), in VHR drone imagery (5cm to
    10cm resolution) however, results might suffer from noise caused by environ-
    ment and drone-based specific problems such as motion artefacts and litter.
    Recent advances in CV based Deep Learning might be able to address these
    issues (Chen et al., 2021 & Carrivick et al., 2016).
    Purpose of the study
    The study is connected to a larger initiative to open-source the assisted
    mapping platform in the current Humanitarian OpenStreetMap (HOTOSM)
    ecosystem. This study is a pilot-project to investigate the capabilities of
    applying semantic segmentation using community open-sourced VHR drone
    imagery collected by the partner organisation OpenAerialMap. The study
    aims to rigourosly assess the various components and inputs that would
    contribute to the ML based mapping system, and to produce a detailed
    evaluation on class-based accuracy assessment (Congalton & Green, 2019).
    This pilot study focuses on 2 camps in East Africa, where data availability
    and the geography of the camps are within a similar savannah ecosystem.
    This enables highly-detailed method testing and analysis of transferability
    of the results between the two camps.
    Data and Methodology
    The first camp is located in Dzaleka, Dowa, Malawi, which is sub-divided
    into the Dzaleka North and Dzaleka main camp. The Dzaleka camps are
    home to around 40,000 refugees mainly coming from the African Great Lakes
    region. The Dzaleka North camp is characterised by a newer, spatially well-
    planned metal-sheeted roofs, while the southern main camp is characterised
    by complex, dense mud-walled building with stone-lined thatched-roofs (UN-
    HCR, 2014). The second camp, the Kalobeyei settlement is part of the sub-
    camp of Kakuma, located in the rural county of Turkana, North-West Kenya.
    The Kalobeyei settlement was home to approximately 34,849 refugees as of
    2019. This camp is significantly more spacious and is characterised by spa-
    tiall well-planned metal-sheeted roofs (UNHCR & DANIDA, 2019). VHR
    drone imageries were provided for both camps and vector labels produced
    by HOTOSM volunteers were provided for the Dzaleka and Dzaleka North
    camp.
    Since CV based Deep Learning is very dependent on the quality of the
    labelled referenced data, especially when performing pixel-based semantic
    segmentaion, it is of crucial importance that care is taken when producing
    highly accurate labels that ensure sucessful training (Ng A., 2018). A large
    quanitiy of available labels did not have such a task in mind, imperfection
    in labelling around existing drone artefacts could cause the trained model
    to misclassify such pixels. In order to train a model which performs well
    on drone imagery, the motion artefact will be a signficant feature for the
    model to learn.he combination of data availability have allow a unique set of
    research questions concerning the input data quality and experiment setup
    to surface. Therefore, to test out U-Net and a few variation of the U-Net
    performance, an additional set of label data was created in order to supple-
    ment the imperfection in the labelled data of the Dzaleka camps. Initially,
    the models will be trained on the pixel-perfect and less complex Kalobeyei
    dataset, this will be then be followed by introducing the Dzaleka datasets
    of higher complexity. A comparison of baseline performance between the U-
    Net variations (Ronneberger et al., 2015) and the Open-Cities-AI-Challenge
    (OCC) winning model is conducted. The baseline experiement aims to keep
    the hyperparameters (e.g. optimiser, learning rate, weight decay etc.) con-
    stant to obtain an objective view of the architectual responses on the same
    dataset setup. This will provide a clear picture of the feasibility and how to
    take this project further, so that further resources could be justified to scale
    future experiments.
    Findings and Discussion
    Initial baseline experiments on the Kalobeyei dataset and Kalobeyei with
    the Dzaleka(s) seem to suggest limited transferability from the OCC model.
    This suggests that the OCC model is perhaps over-generalised to the compe-
    tition test dataset. Despite achieving very high confidence on metal-sheeted
    rooftops, it does not detect any of the more complicated thatched roofs com-
    mon in the Dzaleka camp. The OCC model also struggle with
    some of the more obscure drone motion artefacts occuring at the edge of the
    imagery in the Kalobeyei camp. Meanwhile, the Precision and Recall statis-
    tics favour other variations or further transfer training on the OCC model,
    and the EfficientNet B1 header U-Net pretrained with ImageNet weights.
    However validation loss suggests there might be little room for improvement
    in the further transfer training of the OCC model.
    Precision and Recall have both reached above 0.7 in most experiments,
    which outline the general capability of the strategies used. However there
    are still significant variations among different architectures and setups. The
    next step is to perform systemic fine-tuning to increase the confidence level
    of the appropriate architectures.
    about this event: https://2022.stateofthemap.org/sessions/FRJXCQ/
    6 min
  • Inequalities in the completeness of OpenStreetMap buildings in urban centers (sotm2022)
    Albeit the manifold usage of OSM building footprints an adequate investigation into their completeness on the global scale has not been conducted so far. This talk investigates OSM building completeness within all 13,135 urban centers covering about 50% of the global population.
    The collaborative maps of OpenStreetMap (OSM) have become a major source of geospatial baseline data for humanitarian organisations, companies and public authorities. Describing the elements of spatial data quality (e.g. positional accuracy, completeness, temporal quality) for the OSM dataset is a key prerequisite to provide the potential stakeholders with the necessary information to decide on the fitness for use of a data set for their particular application [1]. Without information on spatial data quality there are serious barriers to the adoption and usage of new sources such as OSM.
    A large community of researchers has analyzed the quality of OSM data in comparison to authoritative reference data sets, by means of remote sensing and using intrinsic measures [2–4]. It has been acknowledged that the OSM data in general is strongly biased, in part due to a much larger contributor basis in countries in the global North as a consequence of socio-economic inequalities and the digital divide [5, 6]. Albeit the manifold usage of OSM building footprints an adequate investigation into their completeness on the global scale has not been conducted so far. This talk investigates OSM building completeness in regions home to a population of 3.5 billion people (about 50% of the global population). First, we propose a machine learning regression method based on generalized additive models (GAMs) to assess OSM building completeness within all 13,135 urban centers (as defined by the European Commission [7]). The analysis utilizes an extensive collection of open building data from commercial and authoritative sources as training data and builds upon very recent technological advances to utilize OSM full-history data for spatio-temporal data analysis on the global scale [8]. This allow us – for the first time – to present a comprehensive assessment of the evolution of urban OSM building completeness which encompasses all data contributed to OSM since 2008.
    For each urban center we calculated the OSM building completeness using the area ratio method which has been applied55 by several other researchers in the context of urban areas [9–11] . Several measures have been adopted to describe the temporal evolution of inequality in urban OSM building mapping on the global scale and per World Bank region. First, we analyzed the share of population living in urban centers with low completeness (<20%) and high completeness (>80%). Gini coefficient has been utilized to derive the degree of evenness of urban OSM building completeness following an approach proposed by Massey & Denton (1988) to study residential segregation 12 . Moran’s I has been selected as a measure of global spatial autocorrelation of urban OSM building completeness. Spatial autocorrelation has been proposed as an explicitly spatial indicator of segregation covering the dimension of clustering [12, 13]. These analyses has been conducted for annual snapshots from 2008-01-01 up until 2022-01-01.
    Overall, urban OSM building completeness is estimated at 38% globally. Our results emphasize that although the well-examined Global North - Global South bias in OSM still exists, over the past years mapping has spread substantially across the globe and within regions. The analysis of the spatial clustering of high completeness values and low completeness values disclosed that global spatial inequality in OSM building completeness has sharply increased between 2008 and 2014. This shows that although overall OSM building completeness became more even in the same period, mapping activity in that time favoured cities which were located close to other cities which were mapped already. One might interprete this as a reinforcing effect. Ongoing mapping in one area triggered even more mapping in surrounded areas. At the same time this also indicates that up until 2014 the expansion of OSM mapping to distant and un-mapped regions (likely to be located in the Global South) didn’t happen at a significant scale.
    Nevertheless, since 2014 Moran’s I global spatial autocorrelation declined and was measured at 0.55 as of 2022. Combined with the decrease of the Gini coefficient in the same time, this suggests that OSM building completeness has become more even because mapping activity has been expanded to regions which were previously mapped much less. In that regard, OSM building data as of today was much less segregated in terms of both dimensions (evenness and clustering) compared to the state-of-the-map in 2014. This process was to a limited extend positively influenced but humanitarian mapping activity organized through the HOT Tasking manager, but hardly influenced by corporate mapping activity.
    We developed a typology of urban centers based on a methodology to quantify intra-urban completeness pattern by means of evenness and spatial clustering. For this we utilized a fine-scale 1x1 km resolution dataset. In total this analysis covered 4,722 urban centers each with a minimum area of 25 square kilometers. Urban centers have been classified into five different types utilizing an agglomerative clustering approach. Our proposed typology of urban centers incorporates the fact that OSM mapping is rarely distributed equal within cities. Similar findings have been reported for Haiti, where densely mapped zones of Port-au-Prince co-exist alongside zones that remain entirely unmapped [14]. Here we provided a method to quantify these pattern and compare across cities.
    The results reveal the need to address the remaining stark data inequalities, which could not be turned around so far by humanitarian and corporate organized mapping activities. We conclude with recommendations directed at stakeholders working with OSM data: (1) Multi-scale building completeness measures should be applied before subsequent usage of OSM data to outline the potential negative effect of missing data. (2) Completeness maps should be used in combination with socio-demographic information to guide future mapping activities to ensure that "nobody is left behind" as encouraged by the SDGs.
    about this event: https://2022.stateofthemap.org/sessions/GPMSLW/
    6 min
  • The cell size issue in OpenStreetMap data quality parameter analyses: an interpolation-based approach (sotm2022)
    The quality of OSM data is dependent on many different factors and is quite heterogeneous. Therefore, in both intrinsic and extrinsic quality parameter analyses, a common practice is subdividing the study areas into subareas. In this paper, we worked on a method for obtaining the optimal grid cell size for OSM data quality analysis. Furthermore, we proposed that if the quality is homogeneous in a region, it can be estimated using an IDW interpolation. . In this summary, we have done a preliminary analysis for a Brazilian city, Curitiba, with about 28,000 points of known accuracy.
    Knowing the quality of a given geospatial data allows measuring how much its use can be viable in specific applications and assist in decision making. ISO 19157 [1] established that the geospatial data quality indicators are positional accuracy, temporal accuracy, thematic accuracy, logical consistency, and completeness. These measures are represented by values that summarize the condition of a product as a whole. These values tend to be homogeneous throughout the evaluated area in traditional mapping. In contrast, in VGI, data quality can be affected by several conditions related to editing history, contribution period, and contributor profiles [8,9]. Given the mentioned aspects, data quality in VGI platforms tends to be heterogeneous, i.e., the results may show significant discrepancies according to the area assessed or even within the same region.
    Given the heterogeneity issues described, several researchers around the world have performed the quality assessment of these types of information based on the principle of subdividing the study area into cells [2,3,4,5,6,7]. Such a procedure has been used in extrinsic quality assessment processes based on ISO 19157 indicators or intrinsic parameters associated with the characteristics of the contributions and contributors. Given the results obtained, the representation of the quality of the data from sub-areas makes it possible to obtain analyses regarding the existence of patterns and establish relationships with other agents and their predominance. The discretization of space into rectangular or hexagonal grids is central to this type of analysis.
    The subdivision can occur regularly or irregularly. The units with irregular dimension cells allow us to perform analyses accepting other features or spatial phenomena that define these dimensions (e.g., neighbourhood border, a river or a railway track, areas with different population densities, and the dichotomy between rural and urban areas). However, these methods make operations difficult because they demand that the area value weigh the values; and the spatial analysis considering the neighbourhood is more complex. Units with regular-sized cells solve these two limitations. However, the problem of the grid of cells not conforming to spatial phenomena or features reappears. In order to conform to them, it is necessary to determine the optimal size of the cells.
    However, one issue remains little discussed: how to determine the size of such cells. Using too large a cell would treat unequal areas equally. On the other hand, using too small a cell and the increased computational cost of the process, ultimately, the ability to generalize the results is lost. Therefore, in this work, we seek to develop an interactive approach for determining the grid cell size calculation, initially using points of known positional accuracy. The hypothesis here is that when the analyzed subarea is of optimal size, one can interpolate the error within the cell via an IDW and generate minimal residuals at the control points. Furthermore, by consecutively subdividing the grid, the mean squared error versus cell size curve will approach stability, thus revealing the optimal size for a given region.
    Inverse distance weighted interpolation (IDW) calculates cell values using sample point sets. This method considers that the higher weights in the interpolation should be due to the proximity of the unknown value point. Thus, if we had a homogeneous behaviour of the quality parameter in an individual area, by interpolation, we could estimate the quality of the points where this value was unknown.
    The methodological procedures developed using python in the QGIS environment are:
    1. For the study area, points of known positional accuracy are chosen (in our case, intersections of the road system), from which a random subset of 10% is separated as a control set.
    2. Definition of a first grid.
    3. The points are used for interpolation within each cell by the IDW method. The Root-mean-square deviation (RMSE) is calculated using the control points for each cell and the average of the RMSEs for the entire area;
    4. Definition of a second grid with half the resolution of the first grid;
    5. Repeat the process described in item 3 for the second grid;
    6. Calculate the differences between the average error values of the second grid and the first grid and check their significance ;
    7. Repetition of the process described in case there are still values indicated as significant.
    In a first analysis, we did a preliminary study for a Brazilian city, Curitiba, with about 28 thousand points of known accuracy. We separated 2.8 thousand control points, and the city was divided into 8 km to 250 m cells. From the preliminary study performed, it was noted that the method show promise in obtaining the necessary analyses to identify the aspects proposed in this work. Furthermore, it was noticed that, as the cell size decreased, the results tended to be more constant, which corroborates the hypothesis of this relationship with data quality. The next steps are to continue the analyses, starting with the verifications and the representation of the magnitude of the differences between different cell sizes.
    Although it is a method that still has a relatively high computational cost to be realized, the results are exciting and can be optimized. It is assumed that if it is possible to identify the minimum cell size in which it is possible to estimate the quality of the features, this will help in decision making regarding the incorporation of procedures in different áreas. This method may need even smaller clippings in regions with very heterogeneous characteristics concerning their surroundings (e.g., slums). It is an initial approach to resolve with data a fundamental issue arising from the lack of knowledge of the granularity of discrepancies for each study area.
    about this event: https://2022.stateofthemap.org/sessions/9HBH3X/
    7 min
  • Exploring Human Bias and Effects of Training in OSM mapping: A Behavioral Experiment in Singapore (sotm2022)
    Human factor is one of the most crucial elements in crowdsourced mapping. This research explores how human bias affects the mapping process and whether such effects can be mitigated through targeted training with a behavioral experiment. The experiment uses a two-group randomized design. The treatment group receives more advanced training than the control group. There are two goals for the experiment. First, we aim to identify the common types of bias that amateurs from a specific demographic community have when using OSM. Second, we plan to explore whether training is helpful for reducing those biases and improve the quality of mapping.
    OpenStreetMap (OSM) is one of the VGI platforms that has been curated primarily by volunteers, which indicates that the demographic differences of the backgrounds of volunteers might affect their understanding of mapping and their mapping behavior.
    For example, contributors with varying skills and experiences of mapping and GIS software might choose different objects to map and trace them in different levels of detail. There is a wide range of factors that could have an impact on how and what individual contributors choose to map: age, gender, expertise, education, income, etc. This study focuses on digging into the mapping behavior of a specific demographic community - residents from Singapore. We have observed changes in terms of tagging and editing behavior in OSM before and after different levels of training. This study has important implications for OSM mapping, especially platforms such as HOTOSM which largely rely on faraway amateur curators to provide up-to-date geographical information of a specific area in case of events such as wars, natural disasters, crimes or humanitarian emergencies.
    about this event: https://2022.stateofthemap.org/sessions/RHF3UX/
    25 min
  • Comparative Integration Potential Analyses of OSM and Wikidata – the Case Study of Railway Stations (sotm2022)
    In this work, we present analyses using a series of comparative data insights that help to better understand the potential and implications of integration between knowledge graphs and OSM.
    OpenStreetMap(OSM) is one of the richest and most diverse sources of geographic information. However, it lacks a fundamental property vital for spatio-semantic analyses: hierarchical structure and semantic linkage. OSM provides links to existing knowledge graphs (structured data that conforms to a specific ontology) e.g., via the wikidata=* tags. The usage of these link-tags is currently limited to a small percentage of both OSM and Wikidata objects. Efforts were undertaken to enhance the geographic linking, linking nearby objects of the same type and semantic linking [1-3]. On the side of the hierarchical and semantic structuring of OSM, the WorldKG knowledge graph[4] provides a semantic mapping of a large subset of OSM. While the free and open OSM tagging scheme is a fundamental part of the OSM project that enabled its success, WorldKG overcomes the inherent lack of structure this tagging scheme represents, paving the way for a knowledge-graph integration of the OSM dataset. Still, open knowledge graphs and OSM are not fully integrated.
    The following analyses provide a series of comparative data insights that help to better understand the potential and implications of integration between knowledge graphs and OSM. In this work, OSM is compared to Wikidata, one of the largest open knowledge graph projects from the Wikimedia Foundation that provides structured storage to other Wikimedia projects such as Wikipedia. Wikidata can, in many aspects, be compared to OSM by its community structure, its free and open nature, and simple contribution framework. In this work, the two datasets are first compared in size, data structure, and distribution. Later, we extend our analyses with a community comparison. The presented analyses also examine how two separate online communities with similar interests have evolved.
    Grasping the size of the two projects is a straightforward task and visible on their websites: OSM features around 1 billion elements [5], while Wikidata is much smaller with over 97 million objects, of which approximately 9 million have geographic coordinates. The topic of railway stations was chosen because these objects have a comparable definition and are well represented in both datasets with ca. 130k and 100k elements in OSM and Wikidata, respectively, indicating integration potential. In OSM, railway stations are mapped by the tags 'railway=station' or 'railway=halt'. In Wikidata, the 'instance of (P31)' property containing 'Q55488' value represents Railway Station (object type).
    By defining generalizable comparison indicators, the presented work provides a framework and source code (available at https://gitlab.gistools.geog.uni-heidelberg.de/giscience/ideal-vgi/osm-wikidata-comparison under the GNU Affero General Public License v3) for VGI project description, comparison, and monitoring. Similar approaches have been established for OSM contributors [6], for single OSM elements [7], and for small geographic regions [8]. For data collection in Wikidata, Wikidata API (https://www.wikidata.org/w/api.php) and Wikidata SPARQL endpoint were used. For Wikidata objects mapped with 'Railway Station', their revision history containing user information, timestamps, and a number of properties was collected. Overall contributions were collected from all users who have contributed to at least one object typed 'Railway Station'. OSM data collection was done using the ohsome API (https://ohsome.org) to extract all railway stations mapped in OSM, including their history and all edits made by the users who edited these railway stations. In addition to a general comparison between the datasets, we derived five sets for a more detailed comparison: OSM with links to Wikidata (59,441 elements), OSM without links to Wikidata (74,659), Wikidata that have links from OSM and are typed as railway stations (45,050), Wikidata without links to OSM but with geocoordinates (54,594) and Wikidata without links to OSM and without geocoordinates (6,714).
    Our first analysis regarding the growth rate of the two sources showed that OSM has reached a saturated state regarding the number of railway stations, where only a few stations were added since mid-2020. Wikidata, on the other hand, still experiences a stable number of new stations that are added to the project. The two datasets depict no clear temporal correlation hinting towards two independent communities, meaning that edits in OSM are not followed by edits in Wikidata and vice versa. Despite the similar size of the two datasets at a global scale, the two datasets show significant discrepancies on a country level. For example, in China, Wikidata features only 39% of the stations present in OSM while having more than double the amount of stations in the United Kingdom. While the lack of stations seems reasonable considering the overall lack of stations in Wikidata, the overabundance of stations in the UK hints towards a data issue that needs more detailed analyses before integration.
    In terms of properties/tags of each object, we observed that Wikidata has, on average more properties per object than OSM. Since Wikidata is a knowledge graph, it also contains links to other objects that can help enrich existing objects increasing this discrepancy even further. OSM objects with links to Wikidatda have almost double the tags compared to those without links. This could either be a quality indicator or an indicator that only famous stations, which are very well mapped in OSM, are also linked to Wikidata. Wikidata objects without links from OSM and geocoordinates have the least number of properties, hinting at their lower quality.
    Next, we present the community analysis. There were around 8.4 million contributors in OSM in total, and 48k unique users have contributed to either creation, deletion, or updating of the railway station objects. In Wikidata, the number of overall contributors is much smaller, i.e., 24k out of which 14k have contributed to Railway objects. The revisions for Wikidata objects are around 11 times higher than that of OSM revisions. This is evident as Wikidata railway stations have more properties than OSM railway stations. This could also be because OSM contributors have a wide variety of objects to map, whether a bench or a tree. In contrast, Wikidata contributors may focus on details and enrichment of prominent objects of public interest. In OSM, adding a new object to the map may take priority over extensive tagging of existing objects. A similar trend is observed for average stations created by each contributor wherein, on average, Wikidata contributors have created five times and, with median statistics, two times more objects than OSM contributors. This may be due to the higher number of bots and imports in Wikidata. While OSM users generally map a specific area that can only feature a limited number of railway stations, Wikidata users may import railway stations from other sources without limiting themselves to a certain geographical unit.
    To conclude, we notice that both communities have great potential to integrate these sources on the topic of railway stations. This potential increases daily with other topics reaching a mature data state in Wikidata and other knowledge graphs. OSM can benefit from the wide range of semantic information linked to objects, while Wikidata can benefit from the precise geoinformation and completeness OSM offers. Yet, care needs to be taken to take both communities on board as each user base exhibits unique data collection styles that need to be respected.
    about this event: https://2022.stateofthemap.org/sessions/YU9JHN/
    27 min

About Chaos Computer Club - archive feed

From the publisher's feed

Der Chaos Computer Club ist die größte europäische Hackervereinigung, und seit über 25 Jahren Vermittler im Spannungsfeld technischer und sozialer Entwicklungen.