Storage Developer Conference

Storage Developer Conference

By SNIA Technical CouncilTechnology
Download on the App Store

Storage Developer Conference episodes

  • #162: Ransomware!!! – an Analysis of Practical Steps for Mitigation and Recovery
    Malware, short for malicious software, is a blanket term for viruses, worms, trojans and other harmful software that attackers use to damage, destroy, and gain access to sensitive information; software is identified as malware based on its intended use, rather than a particular technique or technology used to build it. Ransomware is a blended malware attack that uses a variety of methods to target the victim’s data and then requires the victim to pay a ransom (usually in crypto currency) to the attacker to regain access to the data upon payment (with no guarantees). However, the landscape is changing, and ransomware is no longer just about a financial ransom. Attacks are now being aimed at the infrastructure and undermining public confidence, witness recent headlines regarding incidents affecting police informant databases and oil pipeline sensors. There is also the recent US Treasury guideline to businesses advising them not to pay the ransom. What can we realistically do to prevent such attacks, or do we simply surrender and accept we will lose our data and that the insurance payout will cover any loss? There is increasing evidence that the insurance companies are unwilling to meet those claims, so the situation is perilous as the criminals always appear one step ahead. As a starting point, everyone needs to start assuming they will be attacked at some stage – therefore prevention and mitigation strategies should be based on that assumption. This session outlines the current threats, the scale of the problem, and examines the technology responses currently available as countermeasures. What can be done to prevent an attack? What works and what doesn’t? What should storage developers be thinking about when developing products that need to be more resilient to attack?
    Learning Objectives: 1) Current ransomware trends and scale; 2) Effectiveness of current data protection technology; 3) What other defensive measures should be considered.
    36 min
  • #161: Analysis of Distributed Storage on Blockchain
    Blockchain has revolutionized decentralized finance, and with smart-contracts has enabled the world of Non-Fungible Tokens, set to revolutionize industries such as art, collectibles and gaming. Blockchains, at the very core, are distributed chained hashes. They can be leveraged to store information in a decentralized, secure, encrypted, durable and available format. However, some of the challenges in Blockchain stem from the bloat of storage. Since each participating node will keep a copy of the entire chain, the same data gets replicated on each node, and even a 5MB file stored on the chain can exhaust systems. Several techniques have been used by different implementations to allow Blockchains for distributed storage of data. The advantages compared to cloud storage would be the decentralized nature of storage, the security provided by encrypting content, and the costs. In this session, we will discuss how different blockchain implementations such as Storj, InterPlanetary File System, YottaChain, and ILCOIN have solved the problem of storing data on the chain, but avoiding bloat. Most of these solutions store the data off-chain and store the transactions metadata on the blockchain itself. IPFS & Storj for example, uses content-addressing to uniquely identify each file in a global namespace connecting all the computing devices. The incoming file is encrypted, and split into smaller chunks, and each participating node will store a chunk with zero-knowledge of the other chunks. ILCOIN relies on RIFT protocol to enable two level blockchains, one for the standard blocks, and the other for the mini-blocks that comprise the transactions and which are not mined, but generated by the system. Yottachain uses deduplication after encrypting content, which is not generally the way data storage is designed for cloud, to reduce the footprint of data on the chain. We will discuss the tradeoffs of these solutions and how they aim to disrupt cloud storage. We will compare the benefits provided in terms of security, scalability and costs, and how organizations such as Netflix, Box, Dropbox can benefit from leveraging these technologies.
    Learning Objectives: 1) Learn about the Blockchains and how they are designed to store small amounts of information; 2) Learn about different blockchain projects such as IPFS, Yottachain, Storj, ILCOIN and their implementations; 3) How blockchain based storage solutions can provide better benefits to existing cloud storage; 4) Impact of leveraging blockchain storage on companies such as Netflix, Box, Dropbox.
    24 min
  • #160: SPDK Schedulers
    Polled mode applications such as the Storage Performance Development Kit (SPDK) NVMe over Fabrics target can demonstrate higher performance and efficiency compared to applications with a more traditional interrupt-driven threading model. But this performance and efficiency comes at a cost of increased CPU core utilization when the application is lightly loaded or idle. This talk will introduce a new SPDK scheduler framework which enables transferring work between CPU cores for purposes of shutting down or lowering frequency on cores when under-utilized. We will first describe the SPDK architecture for lightweight threads. Next, we will introduce the scheduler framework and how a scheduler module can collect metrics on the running lightweight threads to make scheduling decisions. Finally, we will share initial results comparing SPDK NVMe-oF target performance and CPU efficiency of a new scheduler module based on this framework with the default static scheduler.
    Learning Objectives: 1) Understand how work is scheduled in an SPDK polled mode application such as the NVMe over Fabrics target; 2) Understand how SPDK scheduler modules under the new framework can decide if and when to move work between CPU cores; 3) Understand how performance + CPU efficiency compare between the default scheduler and scheduler implemented in the new framework.
    33 min
  • #158: NVMe 2.0 Specifications: The Next Generation of NVMe Technology
    The NVM Express® (NVMe®) family of specifications, released in June 2021, allow for faster and simpler development of NVMe solutions to support the increasingly diverse NVMe device environment, now including Hard Disk Drives (HDDs). The extensibility of the specifications encourages the development of independent command sets like Zoned Namespaces (ZNS) and Key Value (KV) while enabling support for the various underlying transport protocols common to NVMe and NVMe over Fabrics (NVMe-oF™) technologies. The NVMe 2.0 library of specifications have been broken out of multiple documents, including the NVMe Base specification, various Command Set specifications, various Transport specifications and the NVMe Management Interface specification. In this session, attendees will learn how the restructured NVMe 2.0 specifications enable the seamless deployment of flash-based solutions in many emerging market segments. This session will provide an overview and usages for several the new features in the NVMe 2.0 Specifications including ZNS, KV, Rotational Media and Endurance Group Management. Finally, the session will cover how these new features will benefit the cloud, enterprise and client market segments.
    Learning Objectives: 1) Learn how the restructured NVMe 2.0 specifications enable the seamless deployment of flash-based solutions in many emerging mark; 2) Gain an overview of usages for several the new features in the NVMe 2.0 Specifications; 3) Learn how the features will benefit the cloud, enterprise and client market segments.
    27 min
  • #157: Compute Express Link 2.0: A High-Performance Interconnect for Memory Pooling
    Data center architectures continue to evolve rapidly to support the ever-growing demands of emerging workloads such as artificial intelligence, machine learning and deep learning. Compute Express Link™ (CXL™) is an open industry-standard interconnect offering coherency and memory semantics using high-bandwidth, low-latency connectivity between the host processor and devices such as accelerators, memory buffers, and smart I/O devices. CXL technology is designed to address the growing needs of high-performance computational workloads by supporting heterogeneous processing and memory systems for applications in artificial intelligence, machine learning, communication systems, and high-performance computing (HPC). These applications deploy a diverse mix of scalar, vector, matrix, and spatial architectures through CPU, GPU, FPGA, smart NICs, and other accelerators. During this session, attendees will learn about the next generation of CXL technology. The CXL 2.0 specification, announced in 2020, adds support for switching for fan-out to connect to more devices; memory pooling for increased memory utilization efficiency and providing memory capacity on demand; and support for persistent memory. This presentation will explore the memory pooling features of CXL 2.0 and how CXL technology will meet the performance and latency demands of emerging workloads for data-hungry applications like AI and ML.
    Learning Objectives: 1)Learn about CXL 2.0, the next generation of Compute Express Link technology; 2) Memory pooling features of CXL 2.0; 3) How CXL will meet the performance and latency demands of emerging workloads for data-hungry applications like AI and ML.
    31 min
  • #156: Quantum Technology and Storage: Where Do They Meet?
    Although quantum technology can be leveraged to do many amazing things, it is not able to provide a general replacement for the storage capabilities we have today with HDDs and SSDs. However, there are a few things where quantum can be leveraged to provide some capabilities that are related to storage and this presentation will cover them. The presentation will start with a quick overview of some of the basic concepts of quantum technology and the reasons why quantum computing may potentially provide significant performance improvements over classical computing for certain applications. It will discuss how quantum computing does implement something similar to computational storage and follow that by explaining how quantum memories can be utilized in certain applications. It will wrap up by explaining how quantum computers work very closely with classical computers to form hybrid classical/quantum processing systems and mention that traditional SSD and HDD storage devices will still be needed on the classical side to support these types of systems.
    40 min
  • #155: Innovations in Load-Store I/O Causing Profound Changes in Memory, Storage, and Compute Landscape
    Emerging and existing applications with cloud computing, 5G, IoT, automotive, and high-performance computing are causing an explosion of data. This data needs to be processed, moved, and stored in a secure, reliable, available, cost-effective, and power-efficient manner. Heterogeneous processing, tiered memory and storage architecture, accelerators, and infrastructure processing units are essential to meet the demands of this evolving compute, memory, and storage landscape. These requirements are driving significant innovations across compute, memory, storage, and interconnect technologies. Compute Express Link* (CXL) with its memory and coherency semantics on top of PCI Express* (PCIe) is paving the way for the convergence of memory and storage with near memory compute capability. Pooling of resources with CXL will lead to rack-scale efficiency with efficient low-latency access mechanisms across multiple nodes in a rack with advanced atomics, acceleration, smart NICs, and persistent memory support. In this talk we will explore how the evolution in load-store interconnects will profoundly change the memory, storage, and compute landscape going forward.
    38 min
  • #154: Amazon FSx For Lustre Deep Dive and its importance in Machine Learning
    Amazon FSx for Lustre, a fully managed service that makes it easy and cost-effective for AWS customers to launch and run a Lustre high performance file system for their data-intensive applications. In this talk I shall introduces you to the features and benefits of the service, such as its massively scalable performance, seamless integration with Amazon S3(object storage), and compatibility with customer applications. Several use cases in particular, we will see how this file system can accelerates and simplifies training Machine Learning models.
    Learning Objectives: Understand Amazon FSx for Lustre internals,How Amazon FSx can be used so that you can accelerate Machine Learning Training,Hands on demo, so that participants can implement the same,Understand how FS Storage is consumed on AWS and design consideration.
    38 min
  • #153: Data Preservation and Retention 101
    This session highlights the differences between retention and preservation. In this session, we will cover: a. Requirements that govern how specific information is maintained b. How long specific information must be kept c. Whether and how specific information is protected and secured d. Understand why data classification is important when setting up data retention policies e. Understand why retention schedules are important f. Understand why regulatory compliance is a major consideration for Preservation requirements
    Learning Objectives: Understand the difference between data preservation and data retention,Understand why data classification is important when setting up data retention policies,Understand why retention schedules are important,Understand why regulatory compliance is a major consideration for Preservation requirements
    56 min

About Storage Developer Conference

From the publisher's feed

Storage developer Podcast, created by developers for developers.