TixelJobs
N
nanvia Indeed

Senior Software Engineer

REMOTEPosted 5mo ago
NLP / LLMSeniorPart-time#python#tensorflow#huggingface#llm#transformers#docker#aws

Not sure if you're a good fit?

Upload your resume and TixelJobs AI will compare it against Senior Software Engineer at nan. Get a match score, missing keywords, and improvement tips before you apply.

Free preview · Your resume stays private

About the Role

# Job Description: Senior Python Developer (Freelance) - High-Performance Web Crawler & AI

## Project Overview

**Scout Crawler** is a high-performance, intelligent web crawler service designed to scrape manufacturer websites for product information at scale. It utilizes a sophisticated **Two-Phase Architecture** (Discovery → Classification → Scraping) to achieve 90% efficiency gains over traditional scrapers. The system handles thousands of concurrent crawls using a distributed Docker-based infrastructure and integrates advanced AI (OpenAI/TensorFlow) for intelligent content filtering.

## The Role

We are seeking an expert **Python Developer** to take ownership of the Scout Crawler's core engine and infrastructure. You will be responsible for maintaining the high-throughput orchestration logic, optimizing resource usage (CPU/Memory) in containerized environments, and ensuring the reliability of the distributed task worker system.

## Key Responsibilities

### 1. Core Architecture & Performance

- **Maintain the Two-Phase Orchestrator**: Manage the complex async logic that drives the `Discovery`, `Classification`, and `Scraping` phases.

- **Optimize Concurrency**: Fine-tune the `asyncio` implementation in `scout-crawler` to maximize throughput without crashing containers (OOM handling).

- **Refine Crawler Strategies**: Improve specific discovery strategies (Sitemap parsing, Heuristic Shallow Crawl, JSON-LD extraction, Platform detection) using `crawl4ai` and `playwright`.

### 2. Infrastructure & Operations

- **Task Management**: oversee the **Task Worker** system, including the auto-recovery logic for orphaned tasks and the `manage_tasks.py` CLI utility.

- **Docker Orchestration**: Manage multi-container deployments (`parallel-task-worker`, `url-discovery`, `api`) and ensure healthy resource reservation/limits in `docker-compose`.

- **Monitoring**: Use the built-in `mass_crawling_monitor` and logging systems to debug stuck crawls, rate limits, or anti-bot blocks.

### 3. AI & Data Pipeline

- **AI Integration**: Manage the `UrlClassificationService`, optimizing OpenAI API usage (cost management) and maintaining fallback to local TensorFlow models.

- **Data Quality**: Ensure accurate extraction of product data and efficient upload of assets to **Digital Ocean Spaces** (S3 compatible).

- **Database Optimization**: Write efficient async SQLAlchemy queries for the PostgreSQL database to handle high-frequency status updates.

## Technical Stack & Requirements

### Must-Have Technical Skills

- **Python 3.10+ & AsyncIO**: Expert-level mastery. You must understand event loops, coroutines, and race conditions deeply.

- **Modern Scraping Stack**: Extensive experience with **Playwright**, **Crawl4AI**, and **BeautifulSoup4**. Knowledge of **Anti-Bot** bypass techniques (stealth cracking, proxies).

- **FastAPI & Pydantic**: Building and maintaining robust REST APIs.

- **Docker & DevOps**: Confident with `docker-compose`, health checks, resource limits, and multi-stage builds.

- **PostgreSQL**: Strong SQL skills, experience with `asyncpg` and `SQLAlchemy 2.0`.

### Nice-to-Have

- **MLOps**: Experience deploying local Transformers (Hugging Face) or managing LLM token budgets.

- **Cloud Infrastructure**: Experience with Digital Ocean App Platform or AWS ECS.

- **System Design**: Understanding of "Controller-Service-Repository" patterns and clean architecture.

## Key Challenges You Will Solve

- **Memory Management**: Preventing memory leaks in long-running headless browser sessions.

- **Distributed Reliability**: Ensuring no task is lost even if a worker crashes, using our recovery mechanisms.

- **Intelligent Filtering**: Tuning the ML classification threshold to balance between precision (getting only products) and recall (not missing hidden items).

## Why This Project?

- **High Impact**: You are not just writing scripts; you are maintaining a high-performance distributed system.

- **Modern Tech**: Work with the latest tools in the Python ecosystem (FastAPI, Pydantic v2, SQLAlchemy 2.0).

- **AI-Native**: Deep integration of Large Language Models into the core scraping workflow.

Job Type: Part-time

Pay: $30.00-$60.00 per hour

Work Location: Remote

Share