Data scientist machine, learning, AWS and data bricks
Data Scientist – Machine Learning, AWS & Databricks
Job Summary
We are seeking an experienced Data Scientist with strong expertise in machine learning, Python, AWS, and Databricks. The ideal candidate will be responsible for analyzing complex datasets, developing predictive models, and deploying scalable machine-learning solutions in cloud environments.
The candidate should have hands-on experience with algorithms such as XGBoost and Random Forest, along with strong proficiency in Python libraries including Pandas and NumPy.
Key Responsibilities
- Collect, clean, transform, and analyze structured and unstructured datasets.
- Perform exploratory data analysis to identify patterns, trends, correlations, and business insights.
- Develop, train, test, and optimize machine-learning models using algorithms such as:
- XGBoost
- Random Forest
- Decision Trees
- Logistic and Linear Regression
- Gradient Boosting
- Clustering and other supervised or unsupervised learning techniques
- Perform feature engineering, feature selection, and hyperparameter tuning.
- Evaluate models using appropriate metrics such as accuracy, precision, recall, F1-score, ROC-AUC, RMSE, and MAE.
- Build reusable data-processing and machine-learning pipelines using Python, Pandas, and NumPy.
- Work with large-scale datasets using Databricks, Apache Spark, and PySpark.
- Develop and deploy cloud-based data science solutions on AWS.
- Build serverless workflows and model-processing components using AWS Lambda.
- Work with AWS services such as S3, SageMaker, Glue, Step Functions, CloudWatch, IAM, and API Gateway.
- Deploy, monitor, maintain, and retrain machine-learning models in production environments.
- Collaborate with data engineers, cloud engineers, software developers, analysts, and business stakeholders.
- Translate business requirements into analytical and machine-learning solutions.
- Document model assumptions, methodologies, performance, limitations, and results.
- Follow machine-learning development, testing, version control, security, and governance best practices.
Required Qualifications
- Bachelor’s or Master’s degree in Computer Science, Data Science, Statistics, Mathematics, Engineering, or a related field.
- Strong programming experience with Python.
- Strong experience with Pandas, NumPy, and Scikit-learn.
- Hands-on experience building machine-learning models using XGBoost and Random Forest.
- Strong understanding of supervised and unsupervised machine-learning techniques.
- Experience with data preprocessing, feature engineering, model validation, and hyperparameter optimization.
- Hands-on experience with Databricks, Spark, or PySpark.
- Experience working with AWS services, particularly AWS Lambda and Amazon S3.
- Strong SQL skills and experience working with relational or cloud-based databases.
- Understanding of model deployment, monitoring, scalability, and production support.
- Experience using Git or another version-control system.
- Strong analytical, problem-solving, and communication skills.
Preferred Qualifications
- Experience with Amazon SageMaker and MLflow.
- Experience developing end-to-end machine-learning pipelines.
- Knowledge of MLOps, CI/CD, Docker, and Infrastructure as Code.
- Experience building REST APIs for machine-learning model inference.
- Knowledge of time-series forecasting, natural language processing, or deep learning.
- Experience working in Agile or Scrum development environments.
- AWS or Databricks certifications are a plus.
Key Technical Skills
Programming: Python, SQL
Data Processing: Pandas, NumPy, PySpark
Machine Learning: XGBoost, Random Forest, Scikit-learn, Gradient Boosting
Cloud: AWS, Lambda, S3, SageMaker, Glue, Step Functions, CloudWatch
Big Data Platform: Databricks, Apache Spark
MLOps and Tools: MLflow, Git, CI/CD, Docker
Pay: $76,735.93 - $92,413.16 per year
Work Location: Hybrid remote in Plano, TX 75023