Research Experience
World Wildlife Fund
- Position: Conservation Data and Technology Intern, Global Science
- Location: San Francisco, CA and New York, NY
Duration: June 2023 – June 2024
- Natural Language Processing specializing in Large Language Models (LLMs)
- Conducted in-depth research on transformer architectures, Large Language Models, focusing on structural innovations.
- Engineered an automated system to process and analyze 500+ unstructured documents using Python and the Unstructured library, enhancing information extraction from diverse formats.
- Spearheaded the development of a Retrieval-Augmented Generation (RAG)-based AI application, integrating multiple LLMs via LangChain to explore new approaches for improved performance and versatility.
- Advanced the field of model optimization by fine-tuning open-source LLMs (e.g., Llama) using Parameter-Efficient Fine-Tuning (PEFT) and Quantized Low-Rank Adaptation (QLoRA) techniques. Achieved a significant 30% accuracy improvement.
- Pioneered advanced prompt engineering methodologies, conducting trend analysis to forecast over 30 project directions to improve human-AI interaction of language models.
- Explored the intersection of LLMs and geospatial technologies, developing novel applications in geocoding and map creation. Utilized LangChain Agent, GeoPandas, Mapbox, Folium, OpenStreetMap, ArcGIS, and Google Search APIs, contributing to the emerging field of AI-enhanced geographic information systems.
- Designed and implemented a Streamlit-based interface for AI applications, focusing on enhancing user interaction with complex language models and improving accessibility for non-technical users.
- Deployed the Generative AI application on Google Cloud Platform, utilizing Docker, App Engine, and Cloud Storage.
- Environmental data lakehouse design
- Conducted comprehensive research on existing environmental data lakes, analyzing their architectures, data models, and performance characteristics. Identified and evaluated suitable designs for large-scale environmental data management, contributing to the field of eco-informatics.
- Conceptualized, designed, and implemented a pilot environmental data lakehouse using Google Cloud technologies (Cloud Storage, BigQuery, Biglake).
- Developed and optimized big data ETL (Extract, Transform, Load) processes for handling and analyzing environmental datasets. Utilized advanced big data technologies including Databricks, PySpark, Hive, and Hadoop.
- Supervisor: Dave Thau
Professional Experience
JD.com, Inc.
- Position: Data Analyst Intern
- Location: Beijing, China
Duration: June 2020 - Aug 2020
- Maintained SQL databases, overseeing data extraction, transformation, and loading (ETL) processes for over 1 million records, ensuring a 99% accuracy rate and reducing processing time by 30%.
- Implemented robust data mining, data processing, and data modeling techniques using Python on a sales dataset with over 50,000 rows, resulting in a 20% increase in predictive accuracy.
- Conducted data analysis of the company’s logistics data using Excel, and delivered actionable insights and strategic recommendations through visualizations with Power BI, resulting in a 15% reduction in operational costs.
