Global ETD Search

Search theses and dissertations gathered from participating repositories worldwide. Every result links back to the library that holds it. No account is needed.

Results

Showing 1 to 11 of 11 for “"text prompts"”.

  1. Quantitative and Qualitative Analysis of Text-to-Image models

    … can create high-quality images from a variety of text prompts. However, a comprehensive analysis that examines both their performance and possible biases is often missing from existing research. In this thesis, I undertake a thorough examination of several leading text-to-image models, namely …

    vt Repository record for Quantitative and Qualitative Analysis of Text-to-Image models (opens in a new tab)

  2. Continual Learning of Object Classification in the Real World

    … and CIFAR100. From the multimodal perspective, text-prompt-based approaches for continual learning leverage pre-trained text encoders and learnable prompts to encode textual features for sequentially arrived classes over time. A common challenge encountered by existing works is how to learn …

    washington Repository record for Continual Learning of Object Classification in the Real World (opens in a new tab)

  3. Text-Guided Image-to-Image Translation for Converting RGB Maps to Tactile Images

    … to generate tactile maps from RGB maps using text-guided image-to-image translation. By leveraging natural language prompts, the method enables control over map details, such as lakes, rivers, and cities, allowing outputs to be tailored to specific needs. A custom dataset of 1,845 RGB maps was …

    carleton Repository record for Text-Guided Image-to-Image Translation for Converting RGB Maps to Tactile Images (opens in a new tab)

  4. Latent diffusion for generative visual attribution in medical image diagnostics

    … generative process, including natural language text prompts acquired from medical science and applied radiology. We perform experiments and quantitatively evaluate our results on the COVID-19 Radiography Database containing labelled chest X-rays with differing pathologies via the Frechet …

    middlesex Repository record for Latent diffusion for generative visual attribution in medical image diagnostics (opens in a new tab)

  5. Multimodal Graphical User Interface for 3D Model Fabrication Through Generative AI

    … of threedimensional assets from natural language prompts and input images, as well as functionalityaware model manipulation through mesh segmentation and categorization. However, all these workflows lack a coherent, unified platform that caters to users’ needs and each method’s technologies. …

    mit Repository record for Multimodal Graphical User Interface for 3D Model Fabrication Through Generative AI (opens in a new tab)

  6. Discovering and Personalizing Artistic Styles with Generative Models

    Text-to-image models have gained widespread popularity, transforming digital art creation by allowing users to generate highly detailed and imaginative visual content from natural language prompts. These models are now widely adopted across various domains, particularly in the arts, where they …

    vt Repository record for Discovering and Personalizing Artistic Styles with Generative Models (opens in a new tab)

  7. The Steerability of Generative Models: Towards Bicycles for the Mind

    … through benchmarks on the steerability of text-to-image and language models. We find that not only is steerability poor, but steering doesn’t reliably improve with more attempts. Third, we propose a framework for designing and optimizing steering mechanisms – tools that help users …

    mit Repository record for The Steerability of Generative Models: Towards Bicycles for the Mind (opens in a new tab)

  8. Bridging the Gap: Generative Machines and Inventive Minds

    … seed media creation, externally expressed in text prompts, sketches, vocalizations, or other intuitive representations. Just as recorded media augmented our ability to perceive and remember, generative media promises to expand our ability to imagine and invent by offering a more immediate path …

    mit Repository record for Bridging the Gap: Generative Machines and Inventive Minds (opens in a new tab)

  9. Multimodal Foundation Models through the Lens of Security: Robust Deepfake Detection and Adversarial Resilience

    … can process multiple types of data, such as text, images, video as input and output, enabling seamless interaction across different modalities. Examples include Text-to-Image (T2I) generation models like DALL-E and Stable Diffusion, which create highly realistic images from simple text

    vt Repository record for Multimodal Foundation Models through the Lens of Security: Robust Deepfake Detection and Adversarial Resilience (opens in a new tab)

  10. Text Prompt-Driven Medical Image Segmentation

    … by introducing rich semantic knowledge from textual descriptions as auxiliary supervision. This integration reduces dependence on manual annotation and computational demand. Although this paradigm has shown impressive success in natural image tasks, its application in medical imaging remains …

    unsw Repository record for Text Prompt-Driven Medical Image Segmentation (opens in a new tab)

  11. Text recaptioning for audio diffusion models

    Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2027-05-01

    uiuc Repository record for Text recaptioning for audio diffusion models (opens in a new tab)