| Nome: | Descrição: | Tamanho: | Formato: | |
|---|---|---|---|---|
| 4.76 MB | Adobe PDF |
Autores
Orientador(es)
Resumo(s)
No contexto académico e profissional o Microsoft PowerPoint é o software mais utilizado
durante a exposição de ideias e apresentações eletrónicas. Contudo, o controlo desta
ferramenta depende fortemente de periféricos tradicionais como o teclado, rato ou comandos
remotos, que por sua vez condicionam a mobilidade, naturalidade do apresentador,
interrompem o fluxo da exposição e implicam custos adicionais de hardware.
A presente dissertação propõe o desenvolvimento de um protótipo funcional com o objetivo
de controlar apresentações PowerPoint recorrendo a reconhecimento de gestos em tempo real,
complementado por comandos de voz, sem a necessidade de qualquer periférico físico
adicional. Tecnologias como computer vision, inteligência artificial e machine learning irão ser
utilizadas paralelemente com camaras convencionais/webcams para detetar e interpretar os
gestos do utilizador, convertendo-os em comandos para o controlo da apresentação.
Adicionalmente, será integrado um modelo de reconhecimento de voz conferindo ao sistema
uma abordagem multimodal, tornando-o mais robusto e resiliente em situações onde o
controlo exclusivo por gestos se revela insuficiente.
Como objetivos específicos, o protótipo pretende incluir a implementação de um módulo de
deteção de gestos com baixa latência e de alta precisão, definindo um conjunto de gestos e
comandos de voz intuitivos para as ações mais comuns, como por exemplo, avançar e recuar a
apresentação. Aditivamente, será realizada a avaliação do protótipo com utilizadores reais,
medindo métricas de precisão, latência e satisfação.
A realização deste trabalho visa contribuir para o avanço tecnológico das Interfaces Naturais de
Utilizador e da interação multimodal, promovendo uma experiência de apresentação fluida,
expressiva e acessível.
In academic and professional contexts, Microsoft PowerPoint is the most widely used software for presenting ideas and electronic presentations. However, control of this tool relies heavily on traditional peripherals such as keyboards, mice or remote controls, which in turn restrict the presenter’s mobility and natural movement, disrupting the flow of the presentation and incurring additional hardware costs. This thesis proposes the development of a functional prototype aimed at controlling PowerPoint presentations using real-time gesture recognition, complemented by voice commands, without the need for any additional physical peripherals. Technologies such as computer vision, artificial intelligence and machine learning will be used in conjunction with conventional webcams to detect and interpret the user’s gestures, converting them into commands to control the presentation. Additionally, a speech recognition model will be integrated, giving the system a multimodal approach and making it more robust and resilient in situations where gesture-only control proves insufficient. As specific objectives, the prototype aims to include the implementation of a low-latency, highprecision gesture detection module, defining a set of standard intuitive gestures and voice commands for the most common actions, such as moving forward and back the presentation. Furthermore, the prototype will be evaluated with real users, measuring metrics of accuracy, latency and satisfaction. This work aims to contribute to the technological advancement of Natural User Interfaces and multimodal interaction, promoting a fluid, expressive and accessible presentation experience.
In academic and professional contexts, Microsoft PowerPoint is the most widely used software for presenting ideas and electronic presentations. However, control of this tool relies heavily on traditional peripherals such as keyboards, mice or remote controls, which in turn restrict the presenter’s mobility and natural movement, disrupting the flow of the presentation and incurring additional hardware costs. This thesis proposes the development of a functional prototype aimed at controlling PowerPoint presentations using real-time gesture recognition, complemented by voice commands, without the need for any additional physical peripherals. Technologies such as computer vision, artificial intelligence and machine learning will be used in conjunction with conventional webcams to detect and interpret the user’s gestures, converting them into commands to control the presentation. Additionally, a speech recognition model will be integrated, giving the system a multimodal approach and making it more robust and resilient in situations where gesture-only control proves insufficient. As specific objectives, the prototype aims to include the implementation of a low-latency, highprecision gesture detection module, defining a set of standard intuitive gestures and voice commands for the most common actions, such as moving forward and back the presentation. Furthermore, the prototype will be evaluated with real users, measuring metrics of accuracy, latency and satisfaction. This work aims to contribute to the technological advancement of Natural User Interfaces and multimodal interaction, promoting a fluid, expressive and accessible presentation experience.
Descrição
Palavras-chave
Interação Humano-computador Interação multimodal Reconhecimento de gestos Comandos de voz Apresentação Microsoft PowerPoint Human-computer interaction Multimodal interaction Gesture recognition Voice commands Microsoft PowerPoint presentation
